Topic recognition - #148

Open
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition
Open

Topic recognition#148
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition

Conversation

@kaia-santosha

Copy link
Copy Markdown

No description provided.

shakes76and others added 27 commits September 17, 2023 21:47
Added Title and overview explanation (description, what problem it solves, etc.)
Added headings for future work on README
All these files are empty but I have created them just so I can go in and edit each one as I progress through the project. My first goal will be to preprocess the dataset in dataset.py
Added a forewarning so markers know why there may be .ipynb files in my future commits
…ataset
I have created a function that returns an adjacency matrix that is created by analysing the edge links depicted in the musae_facebook_edges.csv
Added a normalisation function to get the adjacency matrix in the correct form for the GCN algorithm to work
The features for each node are inconsistent in their quantity, thus I have made a function to convert the feature vector for each node into a bag of word vector so all nodes have equal num of features. If a word is part of a nodes feature list it will be indicated by a 1 and if its not it will be indicated by a 0
since I have seperated the features and labels I have created a train test split for a nodeId list thus can get the train and test features and labels by the ids in the train_ids and test_ids
t-SNE plot created to visualize the initial, high-dimensional node features in a 2D space,
giving insights into their structure and relationships prior to any transformation by the GCN.
This function will be called from train.py in order to access the tensors of the preprocessed data required to be inputted in the model
…ed to research and practicing making my own.
Been a while since last commit, spent researching quite a few projects online for GCN, I attempt to make my own. It is simple for the sake of getting a baseline. I may increase model complexity later on.
needed to import torch and set device for some functionality to work
no hyperparameter tuning yet, just setting some values just to guage whether the model actually trains or not. Hyperparameter tuning will come later when I do cross validation
Created to code to tune hyperparameters via 10-fold nested cross validation. 10 folds were chosen because the dataset is large thus we can afford to train with 90% train and only 10% test sets.
This took an immense time but it was worthwile as I got an indication on the possible best hyperparameters. Since I intend on modifying my model to more align with the gcn code shown in the model exhibition lecture, my hyperparameters may change. However I still think it was worthwile testing the nested cross validation method as I can reuse it to tune the hyperparameters of my final model.
… Exhibition CON session
Though my previous model had good performance, I wanted to experiment and hopefully settle on a new model architecture that better incorporates the techniques learned in the model exhibition CON session
Added the evaluation loop for the model in predict.py, I also added TSNE visualisation after the model is trained to see how well the model has performed. I still have to integrate it properly with train.py by changing some of the train.py code to be modular (in functions) so they can be called by the predict script.
I attempted to commit a local copy of the adjacency_matrix.npy file but since it is 3.9GB it is too large for github. Instead I have added a link to the downloadable file on google drive in the README
fixed some errors in the code namely the return for the train script to the predict script. Also added more detail to the README
wasnt sure which folder should contain the README so Ive put in in both just to be sure
Commented files and double checked if everything is in order
@gayanku

Copy link
Copy Markdown

This is an initial inspection, no action is required at this point
Same as PR #164
Possible duplicate submission?

@shakes76shakes76 added the duplicate This issue or pull request already exists label Nov 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

duplicateThis issue or pull request already existsGCN

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kaia-santosha@gayanku@shakes76@nathasha-naranpanawa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Topic recognition - #148

Open
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition
Open

Topic recognition#148
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition

Conversation

@kaia-santosha

Copy link
Copy Markdown

No description provided.

shakes76and others added 27 commits September 17, 2023 21:47
Added Title and overview explanation (description, what problem it solves, etc.)
Added headings for future work on README
All these files are empty but I have created them just so I can go in and edit each one as I progress through the project. My first goal will be to preprocess the dataset in dataset.py
Added a forewarning so markers know why there may be .ipynb files in my future commits
…ataset
I have created a function that returns an adjacency matrix that is created by analysing the edge links depicted in the musae_facebook_edges.csv
Added a normalisation function to get the adjacency matrix in the correct form for the GCN algorithm to work
The features for each node are inconsistent in their quantity, thus I have made a function to convert the feature vector for each node into a bag of word vector so all nodes have equal num of features. If a word is part of a nodes feature list it will be indicated by a 1 and if its not it will be indicated by a 0
since I have seperated the features and labels I have created a train test split for a nodeId list thus can get the train and test features and labels by the ids in the train_ids and test_ids
t-SNE plot created to visualize the initial, high-dimensional node features in a 2D space,
giving insights into their structure and relationships prior to any transformation by the GCN.
This function will be called from train.py in order to access the tensors of the preprocessed data required to be inputted in the model
…ed to research and practicing making my own.
Been a while since last commit, spent researching quite a few projects online for GCN, I attempt to make my own. It is simple for the sake of getting a baseline. I may increase model complexity later on.
needed to import torch and set device for some functionality to work
no hyperparameter tuning yet, just setting some values just to guage whether the model actually trains or not. Hyperparameter tuning will come later when I do cross validation
Created to code to tune hyperparameters via 10-fold nested cross validation. 10 folds were chosen because the dataset is large thus we can afford to train with 90% train and only 10% test sets.
This took an immense time but it was worthwile as I got an indication on the possible best hyperparameters. Since I intend on modifying my model to more align with the gcn code shown in the model exhibition lecture, my hyperparameters may change. However I still think it was worthwile testing the nested cross validation method as I can reuse it to tune the hyperparameters of my final model.
… Exhibition CON session
Though my previous model had good performance, I wanted to experiment and hopefully settle on a new model architecture that better incorporates the techniques learned in the model exhibition CON session
Added the evaluation loop for the model in predict.py, I also added TSNE visualisation after the model is trained to see how well the model has performed. I still have to integrate it properly with train.py by changing some of the train.py code to be modular (in functions) so they can be called by the predict script.
I attempted to commit a local copy of the adjacency_matrix.npy file but since it is 3.9GB it is too large for github. Instead I have added a link to the downloadable file on google drive in the README
fixed some errors in the code namely the return for the train script to the predict script. Also added more detail to the README
wasnt sure which folder should contain the README so Ive put in in both just to be sure
Commented files and double checked if everything is in order
@gayanku

Copy link
Copy Markdown

This is an initial inspection, no action is required at this point
Same as PR #164
Possible duplicate submission?

@shakes76shakes76 added the duplicate This issue or pull request already exists label Nov 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

duplicateThis issue or pull request already existsGCN

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kaia-santosha@gayanku@shakes76@nathasha-naranpanawa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Topic recognition - #148

Open
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition
Open

Topic recognition#148
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition

Conversation

@kaia-santosha

Copy link
Copy Markdown

No description provided.

shakes76and others added 27 commits September 17, 2023 21:47
Added Title and overview explanation (description, what problem it solves, etc.)
Added headings for future work on README
All these files are empty but I have created them just so I can go in and edit each one as I progress through the project. My first goal will be to preprocess the dataset in dataset.py
Added a forewarning so markers know why there may be .ipynb files in my future commits
…ataset
I have created a function that returns an adjacency matrix that is created by analysing the edge links depicted in the musae_facebook_edges.csv
Added a normalisation function to get the adjacency matrix in the correct form for the GCN algorithm to work
The features for each node are inconsistent in their quantity, thus I have made a function to convert the feature vector for each node into a bag of word vector so all nodes have equal num of features. If a word is part of a nodes feature list it will be indicated by a 1 and if its not it will be indicated by a 0
since I have seperated the features and labels I have created a train test split for a nodeId list thus can get the train and test features and labels by the ids in the train_ids and test_ids
t-SNE plot created to visualize the initial, high-dimensional node features in a 2D space,
giving insights into their structure and relationships prior to any transformation by the GCN.
This function will be called from train.py in order to access the tensors of the preprocessed data required to be inputted in the model
…ed to research and practicing making my own.
Been a while since last commit, spent researching quite a few projects online for GCN, I attempt to make my own. It is simple for the sake of getting a baseline. I may increase model complexity later on.
needed to import torch and set device for some functionality to work
no hyperparameter tuning yet, just setting some values just to guage whether the model actually trains or not. Hyperparameter tuning will come later when I do cross validation
Created to code to tune hyperparameters via 10-fold nested cross validation. 10 folds were chosen because the dataset is large thus we can afford to train with 90% train and only 10% test sets.
This took an immense time but it was worthwile as I got an indication on the possible best hyperparameters. Since I intend on modifying my model to more align with the gcn code shown in the model exhibition lecture, my hyperparameters may change. However I still think it was worthwile testing the nested cross validation method as I can reuse it to tune the hyperparameters of my final model.
… Exhibition CON session
Though my previous model had good performance, I wanted to experiment and hopefully settle on a new model architecture that better incorporates the techniques learned in the model exhibition CON session
Added the evaluation loop for the model in predict.py, I also added TSNE visualisation after the model is trained to see how well the model has performed. I still have to integrate it properly with train.py by changing some of the train.py code to be modular (in functions) so they can be called by the predict script.
I attempted to commit a local copy of the adjacency_matrix.npy file but since it is 3.9GB it is too large for github. Instead I have added a link to the downloadable file on google drive in the README
fixed some errors in the code namely the return for the train script to the predict script. Also added more detail to the README
wasnt sure which folder should contain the README so Ive put in in both just to be sure
Commented files and double checked if everything is in order
@gayanku

Copy link
Copy Markdown

This is an initial inspection, no action is required at this point
Same as PR #164
Possible duplicate submission?

@shakes76shakes76 added the duplicate This issue or pull request already exists label Nov 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

duplicateThis issue or pull request already existsGCN

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kaia-santosha@gayanku@shakes76@nathasha-naranpanawa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Topic recognition - #148

Open
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition
Open

Topic recognition#148
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition

Conversation

@kaia-santosha

Copy link
Copy Markdown

No description provided.

shakes76and others added 27 commits September 17, 2023 21:47
Added Title and overview explanation (description, what problem it solves, etc.)
Added headings for future work on README
All these files are empty but I have created them just so I can go in and edit each one as I progress through the project. My first goal will be to preprocess the dataset in dataset.py
Added a forewarning so markers know why there may be .ipynb files in my future commits
…ataset
I have created a function that returns an adjacency matrix that is created by analysing the edge links depicted in the musae_facebook_edges.csv
Added a normalisation function to get the adjacency matrix in the correct form for the GCN algorithm to work
The features for each node are inconsistent in their quantity, thus I have made a function to convert the feature vector for each node into a bag of word vector so all nodes have equal num of features. If a word is part of a nodes feature list it will be indicated by a 1 and if its not it will be indicated by a 0
since I have seperated the features and labels I have created a train test split for a nodeId list thus can get the train and test features and labels by the ids in the train_ids and test_ids
t-SNE plot created to visualize the initial, high-dimensional node features in a 2D space,
giving insights into their structure and relationships prior to any transformation by the GCN.
This function will be called from train.py in order to access the tensors of the preprocessed data required to be inputted in the model
…ed to research and practicing making my own.
Been a while since last commit, spent researching quite a few projects online for GCN, I attempt to make my own. It is simple for the sake of getting a baseline. I may increase model complexity later on.
needed to import torch and set device for some functionality to work
no hyperparameter tuning yet, just setting some values just to guage whether the model actually trains or not. Hyperparameter tuning will come later when I do cross validation
Created to code to tune hyperparameters via 10-fold nested cross validation. 10 folds were chosen because the dataset is large thus we can afford to train with 90% train and only 10% test sets.
This took an immense time but it was worthwile as I got an indication on the possible best hyperparameters. Since I intend on modifying my model to more align with the gcn code shown in the model exhibition lecture, my hyperparameters may change. However I still think it was worthwile testing the nested cross validation method as I can reuse it to tune the hyperparameters of my final model.
… Exhibition CON session
Though my previous model had good performance, I wanted to experiment and hopefully settle on a new model architecture that better incorporates the techniques learned in the model exhibition CON session
Added the evaluation loop for the model in predict.py, I also added TSNE visualisation after the model is trained to see how well the model has performed. I still have to integrate it properly with train.py by changing some of the train.py code to be modular (in functions) so they can be called by the predict script.
I attempted to commit a local copy of the adjacency_matrix.npy file but since it is 3.9GB it is too large for github. Instead I have added a link to the downloadable file on google drive in the README
fixed some errors in the code namely the return for the train script to the predict script. Also added more detail to the README
wasnt sure which folder should contain the README so Ive put in in both just to be sure
Commented files and double checked if everything is in order
@gayanku

Copy link
Copy Markdown

This is an initial inspection, no action is required at this point
Same as PR #164
Possible duplicate submission?

@shakes76shakes76 added the duplicate This issue or pull request already exists label Nov 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

duplicateThis issue or pull request already existsGCN

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kaia-santosha@gayanku@shakes76@nathasha-naranpanawa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Topic recognition - #148

Open
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition
Open

Topic recognition#148
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition

Conversation

@kaia-santosha

Copy link
Copy Markdown

No description provided.

shakes76and others added 27 commits September 17, 2023 21:47
Added Title and overview explanation (description, what problem it solves, etc.)
Added headings for future work on README
All these files are empty but I have created them just so I can go in and edit each one as I progress through the project. My first goal will be to preprocess the dataset in dataset.py
Added a forewarning so markers know why there may be .ipynb files in my future commits
…ataset
I have created a function that returns an adjacency matrix that is created by analysing the edge links depicted in the musae_facebook_edges.csv
Added a normalisation function to get the adjacency matrix in the correct form for the GCN algorithm to work
The features for each node are inconsistent in their quantity, thus I have made a function to convert the feature vector for each node into a bag of word vector so all nodes have equal num of features. If a word is part of a nodes feature list it will be indicated by a 1 and if its not it will be indicated by a 0
since I have seperated the features and labels I have created a train test split for a nodeId list thus can get the train and test features and labels by the ids in the train_ids and test_ids
t-SNE plot created to visualize the initial, high-dimensional node features in a 2D space,
giving insights into their structure and relationships prior to any transformation by the GCN.
This function will be called from train.py in order to access the tensors of the preprocessed data required to be inputted in the model
…ed to research and practicing making my own.
Been a while since last commit, spent researching quite a few projects online for GCN, I attempt to make my own. It is simple for the sake of getting a baseline. I may increase model complexity later on.
needed to import torch and set device for some functionality to work
no hyperparameter tuning yet, just setting some values just to guage whether the model actually trains or not. Hyperparameter tuning will come later when I do cross validation
Created to code to tune hyperparameters via 10-fold nested cross validation. 10 folds were chosen because the dataset is large thus we can afford to train with 90% train and only 10% test sets.
This took an immense time but it was worthwile as I got an indication on the possible best hyperparameters. Since I intend on modifying my model to more align with the gcn code shown in the model exhibition lecture, my hyperparameters may change. However I still think it was worthwile testing the nested cross validation method as I can reuse it to tune the hyperparameters of my final model.
… Exhibition CON session
Though my previous model had good performance, I wanted to experiment and hopefully settle on a new model architecture that better incorporates the techniques learned in the model exhibition CON session
Added the evaluation loop for the model in predict.py, I also added TSNE visualisation after the model is trained to see how well the model has performed. I still have to integrate it properly with train.py by changing some of the train.py code to be modular (in functions) so they can be called by the predict script.
I attempted to commit a local copy of the adjacency_matrix.npy file but since it is 3.9GB it is too large for github. Instead I have added a link to the downloadable file on google drive in the README
fixed some errors in the code namely the return for the train script to the predict script. Also added more detail to the README
wasnt sure which folder should contain the README so Ive put in in both just to be sure
Commented files and double checked if everything is in order
@gayanku

Copy link
Copy Markdown

This is an initial inspection, no action is required at this point
Same as PR #164
Possible duplicate submission?

@shakes76shakes76 added the duplicate This issue or pull request already exists label Nov 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

duplicateThis issue or pull request already existsGCN

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kaia-santosha@gayanku@shakes76@nathasha-naranpanawa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Topic recognition - #148

Open
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition
Open

Topic recognition#148
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition

Conversation

@kaia-santosha

Copy link
Copy Markdown

No description provided.

shakes76and others added 27 commits September 17, 2023 21:47
Added Title and overview explanation (description, what problem it solves, etc.)
Added headings for future work on README
All these files are empty but I have created them just so I can go in and edit each one as I progress through the project. My first goal will be to preprocess the dataset in dataset.py
Added a forewarning so markers know why there may be .ipynb files in my future commits
…ataset
I have created a function that returns an adjacency matrix that is created by analysing the edge links depicted in the musae_facebook_edges.csv
Added a normalisation function to get the adjacency matrix in the correct form for the GCN algorithm to work
The features for each node are inconsistent in their quantity, thus I have made a function to convert the feature vector for each node into a bag of word vector so all nodes have equal num of features. If a word is part of a nodes feature list it will be indicated by a 1 and if its not it will be indicated by a 0
since I have seperated the features and labels I have created a train test split for a nodeId list thus can get the train and test features and labels by the ids in the train_ids and test_ids
t-SNE plot created to visualize the initial, high-dimensional node features in a 2D space,
giving insights into their structure and relationships prior to any transformation by the GCN.
This function will be called from train.py in order to access the tensors of the preprocessed data required to be inputted in the model
…ed to research and practicing making my own.
Been a while since last commit, spent researching quite a few projects online for GCN, I attempt to make my own. It is simple for the sake of getting a baseline. I may increase model complexity later on.
needed to import torch and set device for some functionality to work
no hyperparameter tuning yet, just setting some values just to guage whether the model actually trains or not. Hyperparameter tuning will come later when I do cross validation
Created to code to tune hyperparameters via 10-fold nested cross validation. 10 folds were chosen because the dataset is large thus we can afford to train with 90% train and only 10% test sets.
This took an immense time but it was worthwile as I got an indication on the possible best hyperparameters. Since I intend on modifying my model to more align with the gcn code shown in the model exhibition lecture, my hyperparameters may change. However I still think it was worthwile testing the nested cross validation method as I can reuse it to tune the hyperparameters of my final model.
… Exhibition CON session
Though my previous model had good performance, I wanted to experiment and hopefully settle on a new model architecture that better incorporates the techniques learned in the model exhibition CON session
Added the evaluation loop for the model in predict.py, I also added TSNE visualisation after the model is trained to see how well the model has performed. I still have to integrate it properly with train.py by changing some of the train.py code to be modular (in functions) so they can be called by the predict script.
I attempted to commit a local copy of the adjacency_matrix.npy file but since it is 3.9GB it is too large for github. Instead I have added a link to the downloadable file on google drive in the README
fixed some errors in the code namely the return for the train script to the predict script. Also added more detail to the README
wasnt sure which folder should contain the README so Ive put in in both just to be sure
Commented files and double checked if everything is in order
@gayanku

Copy link
Copy Markdown

This is an initial inspection, no action is required at this point
Same as PR #164
Possible duplicate submission?

@shakes76shakes76 added the duplicate This issue or pull request already exists label Nov 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

duplicateThis issue or pull request already existsGCN

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kaia-santosha@gayanku@shakes76@nathasha-naranpanawa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Topic recognition - #148

Open
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition
Open

Topic recognition#148
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition

Conversation

@kaia-santosha

Copy link
Copy Markdown

No description provided.

shakes76and others added 27 commits September 17, 2023 21:47
Added Title and overview explanation (description, what problem it solves, etc.)
Added headings for future work on README
All these files are empty but I have created them just so I can go in and edit each one as I progress through the project. My first goal will be to preprocess the dataset in dataset.py
Added a forewarning so markers know why there may be .ipynb files in my future commits
…ataset
I have created a function that returns an adjacency matrix that is created by analysing the edge links depicted in the musae_facebook_edges.csv
Added a normalisation function to get the adjacency matrix in the correct form for the GCN algorithm to work
The features for each node are inconsistent in their quantity, thus I have made a function to convert the feature vector for each node into a bag of word vector so all nodes have equal num of features. If a word is part of a nodes feature list it will be indicated by a 1 and if its not it will be indicated by a 0
since I have seperated the features and labels I have created a train test split for a nodeId list thus can get the train and test features and labels by the ids in the train_ids and test_ids
t-SNE plot created to visualize the initial, high-dimensional node features in a 2D space,
giving insights into their structure and relationships prior to any transformation by the GCN.
This function will be called from train.py in order to access the tensors of the preprocessed data required to be inputted in the model
…ed to research and practicing making my own.
Been a while since last commit, spent researching quite a few projects online for GCN, I attempt to make my own. It is simple for the sake of getting a baseline. I may increase model complexity later on.
needed to import torch and set device for some functionality to work
no hyperparameter tuning yet, just setting some values just to guage whether the model actually trains or not. Hyperparameter tuning will come later when I do cross validation
Created to code to tune hyperparameters via 10-fold nested cross validation. 10 folds were chosen because the dataset is large thus we can afford to train with 90% train and only 10% test sets.
This took an immense time but it was worthwile as I got an indication on the possible best hyperparameters. Since I intend on modifying my model to more align with the gcn code shown in the model exhibition lecture, my hyperparameters may change. However I still think it was worthwile testing the nested cross validation method as I can reuse it to tune the hyperparameters of my final model.
… Exhibition CON session
Though my previous model had good performance, I wanted to experiment and hopefully settle on a new model architecture that better incorporates the techniques learned in the model exhibition CON session
Added the evaluation loop for the model in predict.py, I also added TSNE visualisation after the model is trained to see how well the model has performed. I still have to integrate it properly with train.py by changing some of the train.py code to be modular (in functions) so they can be called by the predict script.
I attempted to commit a local copy of the adjacency_matrix.npy file but since it is 3.9GB it is too large for github. Instead I have added a link to the downloadable file on google drive in the README
fixed some errors in the code namely the return for the train script to the predict script. Also added more detail to the README
wasnt sure which folder should contain the README so Ive put in in both just to be sure
Commented files and double checked if everything is in order
@gayanku

Copy link
Copy Markdown

This is an initial inspection, no action is required at this point
Same as PR #164
Possible duplicate submission?

@shakes76shakes76 added the duplicate This issue or pull request already exists label Nov 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

duplicateThis issue or pull request already existsGCN

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kaia-santosha@gayanku@shakes76@nathasha-naranpanawa
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Topic recognition - #148

Open
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition
Open

Topic recognition#148
kaia-santosha wants to merge 28 commits into
shakes76:mainfrom
kaia-santosha:topic-recognition

Conversation

@kaia-santosha

Copy link
Copy Markdown

No description provided.

shakes76and others added 27 commits September 17, 2023 21:47
Added Title and overview explanation (description, what problem it solves, etc.)
Added headings for future work on README
All these files are empty but I have created them just so I can go in and edit each one as I progress through the project. My first goal will be to preprocess the dataset in dataset.py
Added a forewarning so markers know why there may be .ipynb files in my future commits
…ataset
I have created a function that returns an adjacency matrix that is created by analysing the edge links depicted in the musae_facebook_edges.csv
Added a normalisation function to get the adjacency matrix in the correct form for the GCN algorithm to work
The features for each node are inconsistent in their quantity, thus I have made a function to convert the feature vector for each node into a bag of word vector so all nodes have equal num of features. If a word is part of a nodes feature list it will be indicated by a 1 and if its not it will be indicated by a 0
since I have seperated the features and labels I have created a train test split for a nodeId list thus can get the train and test features and labels by the ids in the train_ids and test_ids
t-SNE plot created to visualize the initial, high-dimensional node features in a 2D space,
giving insights into their structure and relationships prior to any transformation by the GCN.
This function will be called from train.py in order to access the tensors of the preprocessed data required to be inputted in the model
…ed to research and practicing making my own.
Been a while since last commit, spent researching quite a few projects online for GCN, I attempt to make my own. It is simple for the sake of getting a baseline. I may increase model complexity later on.
needed to import torch and set device for some functionality to work
no hyperparameter tuning yet, just setting some values just to guage whether the model actually trains or not. Hyperparameter tuning will come later when I do cross validation
Created to code to tune hyperparameters via 10-fold nested cross validation. 10 folds were chosen because the dataset is large thus we can afford to train with 90% train and only 10% test sets.
This took an immense time but it was worthwile as I got an indication on the possible best hyperparameters. Since I intend on modifying my model to more align with the gcn code shown in the model exhibition lecture, my hyperparameters may change. However I still think it was worthwile testing the nested cross validation method as I can reuse it to tune the hyperparameters of my final model.
… Exhibition CON session
Though my previous model had good performance, I wanted to experiment and hopefully settle on a new model architecture that better incorporates the techniques learned in the model exhibition CON session
Added the evaluation loop for the model in predict.py, I also added TSNE visualisation after the model is trained to see how well the model has performed. I still have to integrate it properly with train.py by changing some of the train.py code to be modular (in functions) so they can be called by the predict script.
I attempted to commit a local copy of the adjacency_matrix.npy file but since it is 3.9GB it is too large for github. Instead I have added a link to the downloadable file on google drive in the README
fixed some errors in the code namely the return for the train script to the predict script. Also added more detail to the README
wasnt sure which folder should contain the README so Ive put in in both just to be sure
Commented files and double checked if everything is in order
@gayanku

Copy link
Copy Markdown

This is an initial inspection, no action is required at this point
Same as PR #164
Possible duplicate submission?

@shakes76shakes76 added the duplicate This issue or pull request already exists label Nov 20, 2023
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

duplicateThis issue or pull request already existsGCN

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@kaia-santosha@gayanku@shakes76@nathasha-naranpanawa