Skip to content

Add arXiv data fetching functionality - #179

Merged
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv
Nov 1, 2025
Merged

Add arXiv data fetching functionality#179
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv

Conversation

@Goziee-git

@Goziee-gitGoziee-git commented Oct 11, 2025

Copy link
Copy Markdown
Contributor

Fixes

Description

Implements comprehensive arXiv data collection system to quantify open access academic papers in the commons.

Type of Change

  • New feature implementing data collection from the arXivopen access academic papers
  • Data source addition/modification

Changes Made

  • Added arXiv API integration for fetching academic paper metadata using the arXiv API as project requirement for automation of fetching new data sources.
  • Implemented data processing pipeline for arXiv submissions in the scripts/1-fetch/arxiv_fetch.py
  • Created filtering logic for open access and CC-licensed papers
  • Added arXiv data to quarterly reporting system

Testing

  • Static analysis passes (./dev/check.sh)
  • arXiv API integration tested with sample queries
  • Data processing validated with test dataset

Data Impact

  • New data source added (arXiv academic papers)
  • Report generation affected (new academic commons metrics)

Related Documentation

  • Updated sources.md with arXiv API credentials setup
  • Added arXiv processing documentation

Checklist

  • I have read and understood the Developer Certificate of Origin (DCO), below, which covers the contents of this pull request (PR).
  • My pull request doesn't include code or content generated with AI.
  • My pull request has a descriptive title (not a vague title like Update index.md).
  • My pull request targets the default branch of the repository (main or master).
  • My commit messages follow best practices.
  • My code follows the established code style of the repository.
  • I added or updated tests for the changes I made (if applicable).
  • I added or updated documentation (if applicable).
  • I tried running the project locally and verified that there are no
    visible errors.

Developer Certificate of Origin

For the purposes of this DCO, "license" is equivalent to "license or public domain dedication," and "open source license" is equivalent to "open content license or public domain dedication."

Developer Certificate of Origin
Developer Certificate of Origin
Version 1.1
Copyright (C) 2004, 2006 The Linux Foundation and its contributors.
1 Letterman Drive
Suite D4700
San Francisco, CA, 94129
Everyone is permitted to copy and distribute verbatim copies of this
license document, but changing it is not allowed.
Developer's Certificate of Origin 1.1
By making a contribution to this project, I certify that:
(a) The contribution was created in whole or in part by me and I
have the right to submit it under the open source license
indicated in the file; or
(b) The contribution is based upon previous work that, to the best
of my knowledge, is covered under an appropriate open source
license and I have the right under that license to submit that
work with modifications, whether created in whole or in part
by me, under the same open source license (unless I am
permitted to submit under a different license), as indicated
in the file; or
(c) The contribution was provided directly to me by some other
person who certified (a), (b) or (c) and I have not modified
it.
(d) I understand and agree that this project and the contribution
are public and that a record of the contribution (including all
personal information I submit with it, including my sign-off) is
maintained indefinitely and may be redistributed consistent with
this project or the open source license(s) involved.

@Goziee-git
Goziee-git requested review from a team as code ownersOctober 11, 2025 23:29
@Goziee-git
Goziee-git requested review from TimidRobot and possumbilities and removed request for a teamOctober 11, 2025 23:29
@cc-open-source-botcc-open-source-bot moved this to In review in TimidRobotOct 11, 2025
@Goziee-gitGoziee-git changed the title Add arXiv data fetching and processing functionalityAdd arXiv data fetching functionalityOct 12, 2025

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

Comment threadscripts/3-report/gcs_report.py
Comment threaddata/2025Q4/1-fetch/arxiv_1_count.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_category.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_language.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_country.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_year.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_4_count_by_author_count.csv Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

This comment was marked as outdated.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please follow through on your first pull request (PR) before submitting any more:

Depending on how that one goes, I might reopen this PR.

hello @TimidRobot as requested, I have made changes based on your review. Good work is emphasized over speed and I do hope my attempt to go full circle with other PR hasn't dented my chances of significantly contributing to the project. Thank you🙏🏼

@TimidRobotTimidRobot reopened this Oct 17, 2025
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git ok, please focus on this PR

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

@TimidRobot I have removed the query for languages as it returns only English. Also worthy of note here, the arxiv data source accepts papers in other languages but requires that the paper abstracts be submitted in English. so Impossible to get a good distribution of licenses as per language usage.

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time. The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+). The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

Comment threadscripts/3-report/gcs_report.py
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py
Comment threadarxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time.

  • How was arxiv_category_map.yml created?
    • If by script, it should probably go in dev/
  • Data that persists should go in data/ not a specific quarter directory

The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+).

I'll look at data after outstanding comments are resolved.

The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

I'll look at data after outstanding comments are resolved. That said, I'm not excited about adding JSON to the project.

@TimidRobot

Copy link
Copy Markdown
Member

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

It's probably a good idea to create a function in the shared library eventually. Please leave that to last, however.

Refactor arxiv_fetch.py to use requests library for HTTP requests, implementing retry logic for better error handling. Update license extraction logic and CSV headers to remove PLAN_INDEX.
@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot Hello, are there any more changes you'll like to be made to the ones already identified and done. Also i'ld like to ask, if there are no more changes to be made on this, what steps would you recommend as the next one for me in the project, with respect to integrating ArXiv as a valid source for the project?

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Please update sources.md for arXiv

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git Please update sources.md for arXiv

@TimidRobot, Thanks, Im happy we've gotten to this point. very excited to continue contributing to the project. I would like your permission to continue working on the processing and reporting scripts for the arXiv data source

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot as per #179 (comment)
the implementation of the bucketing for the AUTHOR_COUNT was because you suggested that it was not so insightful to just merely dump the data, and so i thought a sort of grouping of this data would help us to generate meaningful insights. Raw author counts would create extremely sparse data with many single-occurrence values while Bucketing ("1", "2-3", "4-6", "7-10", "11+") on the other hand, creates meaningful statistical groups for analysis. It also facilitates visualization and reporting by creating compact, aggregated datasets that are easier to process and analyze hence supporting the project's goal of understanding "how knowledge and culture of the commons is distributed".

although these are my thoughts, I am still very happy to know what you think about this. To further share my opinion, these are some helpful links that guided my decision for bucketing the AUTHOR_COUNT data.

Scientometrics Journal Guidelines: (https://link.springer.com/journal/11192)
• Standard practice in bibliometric studies to group author counts into ranges for collaboration analysis which reduces noise and enables meaningful statistical comparisons across disciplines

FAIR Data Principles: (https://www.go-fair.org/fair-principles/)
• Bucketing supports Findability and Reusability by creating standardized categorical data
NIH Data Management Guidelines: (https://sharing.nih.gov/data-management-and-sharing-policy)

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please see two unresolved conversations:

  1. Add arXiv data fetching functionality #179 (comment)
  2. Add arXiv data fetching functionality #179 (comment)

@TimidRobot the conversations mentioned here are identical, please i'ld like to know what your preference is with the AUTHOR_COUNT, especially if you want it removed or modified in differnt way

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Sorry, let me clarify. I think bucketing authors into AUTHOR_BUCKET is a good and helpful course of action. I wouldd also add https://en.wikipedia.org/wiki/Data_binning to the list of references.

I think which buckets are selected could use some refinement. My thinking is that there are three considerations:

  1. plotting
  2. precedence
  3. distribution

Plotting

For plotting, I think around five values display well.

Precedence

For precedence, we can look at the various citation styles:

Distribution

The script can be modified to provide information on all relevant author counts (only results above 1% are shown):

AuthorsCountPercent
313122.55%
211119.10%
49215.83%
18514.63%
5467.92%
6233.96%
7213.61%
8193.27%
9152.58%
10101.72%

Recommendation

I recommend the following buckets:

  • 1 author
  • 2 authors
  • 3 authors
  • 4 authors
  • 5+ authors

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@Goziee-git

Goziee-git commented Oct 31, 2025

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot as per ordering the constants, what would be your preference here. The current order of the constants is:

  1. FILE_ARXIV_COUNT (arxiv_1_count.csv)
  2. FILE_ARXIV_CATEGORY_REPORT (arxiv_2_count_by_category_report.csv)
  3. FILE_ARXIV_YEAR (arxiv_3_count_by_year.csv)
  4. FILE_ARXIV_AUTHOR_BUCKET (arxiv_4_count_by_author_bucket.csv)

This follows the logical/sequential order based on the numbered workflow (1 → 2 → 3 → 4).

in comparison to the gcs_fetch.py (current order):

  1. FILE1_COUNT (gcs_1_count.csv)
  2. FILE2_LANGUAGE (gcs_2_count_by_language.csv)
  3. FILE3_COUNTRY (gcs_3_count_by_country.csv)

Comparison:

  • Both scripts follow the same logical/sequential ordering pattern
  • Both order constants by their numbered workflow sequence (1 → 2 → 3 → 4)
  • Both use the numbering in the filename to determine order

Do you prefer some other order for this, so i can be guided as to the best course of action here

@Babi-B

Babi-B commented Oct 31, 2025

Copy link
Copy Markdown
Contributor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before
FILE_ARXIV_CATEGORY_REPORT

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before FILE_ARXIV_CATEGORY_REPORT

@Babi-B, Thanks for the clarity, i appreciate your reviews here.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Babi-B, as per #185 looks like you've not made additions to sources.md. A moment ago i tried to read through and no links or references. I think you should update that. Or would you like me to go ahead with that 🚀🚀🚀🚀🚀

@TimidRobot

Copy link
Copy Markdown
Member

Both use the numbering in the filename to determine order

@Goziee-git The difference is that gcs_fetch.py uses a number in the constant so that they follow the flow when ordered:

FILE1_COUNT
^
FILE2_LANGUAGE
^

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work, thank you!!

@TimidRobot
TimidRobot merged commit 0d44547 into creativecommons:mainNov 1, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Archived in project

Development

Successfully merging this pull request may close these issues.

Integrate arXiv as data source for academic commons quantification

4 participants

@Goziee-git@TimidRobot@Babi-B@cc-open-source-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Add arXiv data fetching functionality by Goziee-git · Pull Request #179 · creativecommons/quantifying · GitHub
Skip to content

Add arXiv data fetching functionality - #179

Merged
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv
Nov 1, 2025
Merged

Add arXiv data fetching functionality#179
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv

Conversation

@Goziee-git

@Goziee-gitGoziee-git commented Oct 11, 2025

Copy link
Copy Markdown
Contributor

Fixes

Description

Implements comprehensive arXiv data collection system to quantify open access academic papers in the commons.

Type of Change

  • New feature implementing data collection from the arXivopen access academic papers
  • Data source addition/modification

Changes Made

  • Added arXiv API integration for fetching academic paper metadata using the arXiv API as project requirement for automation of fetching new data sources.
  • Implemented data processing pipeline for arXiv submissions in the scripts/1-fetch/arxiv_fetch.py
  • Created filtering logic for open access and CC-licensed papers
  • Added arXiv data to quarterly reporting system

Testing

  • Static analysis passes (./dev/check.sh)
  • arXiv API integration tested with sample queries
  • Data processing validated with test dataset

Data Impact

  • New data source added (arXiv academic papers)
  • Report generation affected (new academic commons metrics)

Related Documentation

  • Updated sources.md with arXiv API credentials setup
  • Added arXiv processing documentation

Checklist

  • I have read and understood the Developer Certificate of Origin (DCO), below, which covers the contents of this pull request (PR).
  • My pull request doesn't include code or content generated with AI.
  • My pull request has a descriptive title (not a vague title like Update index.md).
  • My pull request targets the default branch of the repository (main or master).
  • My commit messages follow best practices.
  • My code follows the established code style of the repository.
  • I added or updated tests for the changes I made (if applicable).
  • I added or updated documentation (if applicable).
  • I tried running the project locally and verified that there are no
    visible errors.

Developer Certificate of Origin

For the purposes of this DCO, "license" is equivalent to "license or public domain dedication," and "open source license" is equivalent to "open content license or public domain dedication."

Developer Certificate of Origin
Developer Certificate of Origin
Version 1.1
Copyright (C) 2004, 2006 The Linux Foundation and its contributors.
1 Letterman Drive
Suite D4700
San Francisco, CA, 94129
Everyone is permitted to copy and distribute verbatim copies of this
license document, but changing it is not allowed.
Developer's Certificate of Origin 1.1
By making a contribution to this project, I certify that:
(a) The contribution was created in whole or in part by me and I
have the right to submit it under the open source license
indicated in the file; or
(b) The contribution is based upon previous work that, to the best
of my knowledge, is covered under an appropriate open source
license and I have the right under that license to submit that
work with modifications, whether created in whole or in part
by me, under the same open source license (unless I am
permitted to submit under a different license), as indicated
in the file; or
(c) The contribution was provided directly to me by some other
person who certified (a), (b) or (c) and I have not modified
it.
(d) I understand and agree that this project and the contribution
are public and that a record of the contribution (including all
personal information I submit with it, including my sign-off) is
maintained indefinitely and may be redistributed consistent with
this project or the open source license(s) involved.

@Goziee-git
Goziee-git requested review from a team as code ownersOctober 11, 2025 23:29
@Goziee-git
Goziee-git requested review from TimidRobot and possumbilities and removed request for a teamOctober 11, 2025 23:29
@cc-open-source-botcc-open-source-bot moved this to In review in TimidRobotOct 11, 2025
@Goziee-gitGoziee-git changed the title Add arXiv data fetching and processing functionalityAdd arXiv data fetching functionalityOct 12, 2025

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

Comment threadscripts/3-report/gcs_report.py
Comment threaddata/2025Q4/1-fetch/arxiv_1_count.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_category.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_language.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_country.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_year.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_4_count_by_author_count.csv Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

This comment was marked as outdated.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please follow through on your first pull request (PR) before submitting any more:

Depending on how that one goes, I might reopen this PR.

hello @TimidRobot as requested, I have made changes based on your review. Good work is emphasized over speed and I do hope my attempt to go full circle with other PR hasn't dented my chances of significantly contributing to the project. Thank you🙏🏼

@TimidRobotTimidRobot reopened this Oct 17, 2025
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git ok, please focus on this PR

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

@TimidRobot I have removed the query for languages as it returns only English. Also worthy of note here, the arxiv data source accepts papers in other languages but requires that the paper abstracts be submitted in English. so Impossible to get a good distribution of licenses as per language usage.

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time. The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+). The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

Comment threadscripts/3-report/gcs_report.py
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py
Comment threadarxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time.

  • How was arxiv_category_map.yml created?
    • If by script, it should probably go in dev/
  • Data that persists should go in data/ not a specific quarter directory

The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+).

I'll look at data after outstanding comments are resolved.

The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

I'll look at data after outstanding comments are resolved. That said, I'm not excited about adding JSON to the project.

@TimidRobot

Copy link
Copy Markdown
Member

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

It's probably a good idea to create a function in the shared library eventually. Please leave that to last, however.

Refactor arxiv_fetch.py to use requests library for HTTP requests, implementing retry logic for better error handling. Update license extraction logic and CSV headers to remove PLAN_INDEX.
@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot Hello, are there any more changes you'll like to be made to the ones already identified and done. Also i'ld like to ask, if there are no more changes to be made on this, what steps would you recommend as the next one for me in the project, with respect to integrating ArXiv as a valid source for the project?

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Please update sources.md for arXiv

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git Please update sources.md for arXiv

@TimidRobot, Thanks, Im happy we've gotten to this point. very excited to continue contributing to the project. I would like your permission to continue working on the processing and reporting scripts for the arXiv data source

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot as per #179 (comment)
the implementation of the bucketing for the AUTHOR_COUNT was because you suggested that it was not so insightful to just merely dump the data, and so i thought a sort of grouping of this data would help us to generate meaningful insights. Raw author counts would create extremely sparse data with many single-occurrence values while Bucketing ("1", "2-3", "4-6", "7-10", "11+") on the other hand, creates meaningful statistical groups for analysis. It also facilitates visualization and reporting by creating compact, aggregated datasets that are easier to process and analyze hence supporting the project's goal of understanding "how knowledge and culture of the commons is distributed".

although these are my thoughts, I am still very happy to know what you think about this. To further share my opinion, these are some helpful links that guided my decision for bucketing the AUTHOR_COUNT data.

Scientometrics Journal Guidelines: (https://link.springer.com/journal/11192)
• Standard practice in bibliometric studies to group author counts into ranges for collaboration analysis which reduces noise and enables meaningful statistical comparisons across disciplines

FAIR Data Principles: (https://www.go-fair.org/fair-principles/)
• Bucketing supports Findability and Reusability by creating standardized categorical data
NIH Data Management Guidelines: (https://sharing.nih.gov/data-management-and-sharing-policy)

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please see two unresolved conversations:

  1. Add arXiv data fetching functionality #179 (comment)
  2. Add arXiv data fetching functionality #179 (comment)

@TimidRobot the conversations mentioned here are identical, please i'ld like to know what your preference is with the AUTHOR_COUNT, especially if you want it removed or modified in differnt way

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Sorry, let me clarify. I think bucketing authors into AUTHOR_BUCKET is a good and helpful course of action. I wouldd also add https://en.wikipedia.org/wiki/Data_binning to the list of references.

I think which buckets are selected could use some refinement. My thinking is that there are three considerations:

  1. plotting
  2. precedence
  3. distribution

Plotting

For plotting, I think around five values display well.

Precedence

For precedence, we can look at the various citation styles:

Distribution

The script can be modified to provide information on all relevant author counts (only results above 1% are shown):

AuthorsCountPercent
313122.55%
211119.10%
49215.83%
18514.63%
5467.92%
6233.96%
7213.61%
8193.27%
9152.58%
10101.72%

Recommendation

I recommend the following buckets:

  • 1 author
  • 2 authors
  • 3 authors
  • 4 authors
  • 5+ authors

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@Goziee-git

Goziee-git commented Oct 31, 2025

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot as per ordering the constants, what would be your preference here. The current order of the constants is:

  1. FILE_ARXIV_COUNT (arxiv_1_count.csv)
  2. FILE_ARXIV_CATEGORY_REPORT (arxiv_2_count_by_category_report.csv)
  3. FILE_ARXIV_YEAR (arxiv_3_count_by_year.csv)
  4. FILE_ARXIV_AUTHOR_BUCKET (arxiv_4_count_by_author_bucket.csv)

This follows the logical/sequential order based on the numbered workflow (1 → 2 → 3 → 4).

in comparison to the gcs_fetch.py (current order):

  1. FILE1_COUNT (gcs_1_count.csv)
  2. FILE2_LANGUAGE (gcs_2_count_by_language.csv)
  3. FILE3_COUNTRY (gcs_3_count_by_country.csv)

Comparison:

  • Both scripts follow the same logical/sequential ordering pattern
  • Both order constants by their numbered workflow sequence (1 → 2 → 3 → 4)
  • Both use the numbering in the filename to determine order

Do you prefer some other order for this, so i can be guided as to the best course of action here

@Babi-B

Babi-B commented Oct 31, 2025

Copy link
Copy Markdown
Contributor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before
FILE_ARXIV_CATEGORY_REPORT

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before FILE_ARXIV_CATEGORY_REPORT

@Babi-B, Thanks for the clarity, i appreciate your reviews here.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Babi-B, as per #185 looks like you've not made additions to sources.md. A moment ago i tried to read through and no links or references. I think you should update that. Or would you like me to go ahead with that 🚀🚀🚀🚀🚀

@TimidRobot

Copy link
Copy Markdown
Member

Both use the numbering in the filename to determine order

@Goziee-git The difference is that gcs_fetch.py uses a number in the constant so that they follow the flow when ordered:

FILE1_COUNT
^
FILE2_LANGUAGE
^

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work, thank you!!

@TimidRobot
TimidRobot merged commit 0d44547 into creativecommons:mainNov 1, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Archived in project

Development

Successfully merging this pull request may close these issues.

Integrate arXiv as data source for academic commons quantification

4 participants

@Goziee-git@TimidRobot@Babi-B@cc-open-source-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add arXiv data fetching functionality by Goziee-git · Pull Request #179 · creativecommons/quantifying · GitHub
Skip to content

Add arXiv data fetching functionality - #179

Merged
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv
Nov 1, 2025
Merged

Add arXiv data fetching functionality#179
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv

Conversation

@Goziee-git

@Goziee-gitGoziee-git commented Oct 11, 2025

Copy link
Copy Markdown
Contributor

Fixes

Description

Implements comprehensive arXiv data collection system to quantify open access academic papers in the commons.

Type of Change

  • New feature implementing data collection from the arXivopen access academic papers
  • Data source addition/modification

Changes Made

  • Added arXiv API integration for fetching academic paper metadata using the arXiv API as project requirement for automation of fetching new data sources.
  • Implemented data processing pipeline for arXiv submissions in the scripts/1-fetch/arxiv_fetch.py
  • Created filtering logic for open access and CC-licensed papers
  • Added arXiv data to quarterly reporting system

Testing

  • Static analysis passes (./dev/check.sh)
  • arXiv API integration tested with sample queries
  • Data processing validated with test dataset

Data Impact

  • New data source added (arXiv academic papers)
  • Report generation affected (new academic commons metrics)

Related Documentation

  • Updated sources.md with arXiv API credentials setup
  • Added arXiv processing documentation

Checklist

  • I have read and understood the Developer Certificate of Origin (DCO), below, which covers the contents of this pull request (PR).
  • My pull request doesn't include code or content generated with AI.
  • My pull request has a descriptive title (not a vague title like Update index.md).
  • My pull request targets the default branch of the repository (main or master).
  • My commit messages follow best practices.
  • My code follows the established code style of the repository.
  • I added or updated tests for the changes I made (if applicable).
  • I added or updated documentation (if applicable).
  • I tried running the project locally and verified that there are no
    visible errors.

Developer Certificate of Origin

For the purposes of this DCO, "license" is equivalent to "license or public domain dedication," and "open source license" is equivalent to "open content license or public domain dedication."

Developer Certificate of Origin
Developer Certificate of Origin
Version 1.1
Copyright (C) 2004, 2006 The Linux Foundation and its contributors.
1 Letterman Drive
Suite D4700
San Francisco, CA, 94129
Everyone is permitted to copy and distribute verbatim copies of this
license document, but changing it is not allowed.
Developer's Certificate of Origin 1.1
By making a contribution to this project, I certify that:
(a) The contribution was created in whole or in part by me and I
have the right to submit it under the open source license
indicated in the file; or
(b) The contribution is based upon previous work that, to the best
of my knowledge, is covered under an appropriate open source
license and I have the right under that license to submit that
work with modifications, whether created in whole or in part
by me, under the same open source license (unless I am
permitted to submit under a different license), as indicated
in the file; or
(c) The contribution was provided directly to me by some other
person who certified (a), (b) or (c) and I have not modified
it.
(d) I understand and agree that this project and the contribution
are public and that a record of the contribution (including all
personal information I submit with it, including my sign-off) is
maintained indefinitely and may be redistributed consistent with
this project or the open source license(s) involved.

@Goziee-git
Goziee-git requested review from a team as code ownersOctober 11, 2025 23:29
@Goziee-git
Goziee-git requested review from TimidRobot and possumbilities and removed request for a teamOctober 11, 2025 23:29
@cc-open-source-botcc-open-source-bot moved this to In review in TimidRobotOct 11, 2025
@Goziee-gitGoziee-git changed the title Add arXiv data fetching and processing functionalityAdd arXiv data fetching functionalityOct 12, 2025

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

Comment threadscripts/3-report/gcs_report.py
Comment threaddata/2025Q4/1-fetch/arxiv_1_count.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_category.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_language.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_country.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_year.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_4_count_by_author_count.csv Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

This comment was marked as outdated.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please follow through on your first pull request (PR) before submitting any more:

Depending on how that one goes, I might reopen this PR.

hello @TimidRobot as requested, I have made changes based on your review. Good work is emphasized over speed and I do hope my attempt to go full circle with other PR hasn't dented my chances of significantly contributing to the project. Thank you🙏🏼

@TimidRobotTimidRobot reopened this Oct 17, 2025
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git ok, please focus on this PR

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

@TimidRobot I have removed the query for languages as it returns only English. Also worthy of note here, the arxiv data source accepts papers in other languages but requires that the paper abstracts be submitted in English. so Impossible to get a good distribution of licenses as per language usage.

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time. The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+). The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

Comment threadscripts/3-report/gcs_report.py
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py
Comment threadarxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time.

  • How was arxiv_category_map.yml created?
    • If by script, it should probably go in dev/
  • Data that persists should go in data/ not a specific quarter directory

The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+).

I'll look at data after outstanding comments are resolved.

The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

I'll look at data after outstanding comments are resolved. That said, I'm not excited about adding JSON to the project.

@TimidRobot

Copy link
Copy Markdown
Member

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

It's probably a good idea to create a function in the shared library eventually. Please leave that to last, however.

Refactor arxiv_fetch.py to use requests library for HTTP requests, implementing retry logic for better error handling. Update license extraction logic and CSV headers to remove PLAN_INDEX.
@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot Hello, are there any more changes you'll like to be made to the ones already identified and done. Also i'ld like to ask, if there are no more changes to be made on this, what steps would you recommend as the next one for me in the project, with respect to integrating ArXiv as a valid source for the project?

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Please update sources.md for arXiv

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git Please update sources.md for arXiv

@TimidRobot, Thanks, Im happy we've gotten to this point. very excited to continue contributing to the project. I would like your permission to continue working on the processing and reporting scripts for the arXiv data source

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot as per #179 (comment)
the implementation of the bucketing for the AUTHOR_COUNT was because you suggested that it was not so insightful to just merely dump the data, and so i thought a sort of grouping of this data would help us to generate meaningful insights. Raw author counts would create extremely sparse data with many single-occurrence values while Bucketing ("1", "2-3", "4-6", "7-10", "11+") on the other hand, creates meaningful statistical groups for analysis. It also facilitates visualization and reporting by creating compact, aggregated datasets that are easier to process and analyze hence supporting the project's goal of understanding "how knowledge and culture of the commons is distributed".

although these are my thoughts, I am still very happy to know what you think about this. To further share my opinion, these are some helpful links that guided my decision for bucketing the AUTHOR_COUNT data.

Scientometrics Journal Guidelines: (https://link.springer.com/journal/11192)
• Standard practice in bibliometric studies to group author counts into ranges for collaboration analysis which reduces noise and enables meaningful statistical comparisons across disciplines

FAIR Data Principles: (https://www.go-fair.org/fair-principles/)
• Bucketing supports Findability and Reusability by creating standardized categorical data
NIH Data Management Guidelines: (https://sharing.nih.gov/data-management-and-sharing-policy)

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please see two unresolved conversations:

  1. Add arXiv data fetching functionality #179 (comment)
  2. Add arXiv data fetching functionality #179 (comment)

@TimidRobot the conversations mentioned here are identical, please i'ld like to know what your preference is with the AUTHOR_COUNT, especially if you want it removed or modified in differnt way

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Sorry, let me clarify. I think bucketing authors into AUTHOR_BUCKET is a good and helpful course of action. I wouldd also add https://en.wikipedia.org/wiki/Data_binning to the list of references.

I think which buckets are selected could use some refinement. My thinking is that there are three considerations:

  1. plotting
  2. precedence
  3. distribution

Plotting

For plotting, I think around five values display well.

Precedence

For precedence, we can look at the various citation styles:

Distribution

The script can be modified to provide information on all relevant author counts (only results above 1% are shown):

AuthorsCountPercent
313122.55%
211119.10%
49215.83%
18514.63%
5467.92%
6233.96%
7213.61%
8193.27%
9152.58%
10101.72%

Recommendation

I recommend the following buckets:

  • 1 author
  • 2 authors
  • 3 authors
  • 4 authors
  • 5+ authors

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@Goziee-git

Goziee-git commented Oct 31, 2025

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot as per ordering the constants, what would be your preference here. The current order of the constants is:

  1. FILE_ARXIV_COUNT (arxiv_1_count.csv)
  2. FILE_ARXIV_CATEGORY_REPORT (arxiv_2_count_by_category_report.csv)
  3. FILE_ARXIV_YEAR (arxiv_3_count_by_year.csv)
  4. FILE_ARXIV_AUTHOR_BUCKET (arxiv_4_count_by_author_bucket.csv)

This follows the logical/sequential order based on the numbered workflow (1 → 2 → 3 → 4).

in comparison to the gcs_fetch.py (current order):

  1. FILE1_COUNT (gcs_1_count.csv)
  2. FILE2_LANGUAGE (gcs_2_count_by_language.csv)
  3. FILE3_COUNTRY (gcs_3_count_by_country.csv)

Comparison:

  • Both scripts follow the same logical/sequential ordering pattern
  • Both order constants by their numbered workflow sequence (1 → 2 → 3 → 4)
  • Both use the numbering in the filename to determine order

Do you prefer some other order for this, so i can be guided as to the best course of action here

@Babi-B

Babi-B commented Oct 31, 2025

Copy link
Copy Markdown
Contributor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before
FILE_ARXIV_CATEGORY_REPORT

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before FILE_ARXIV_CATEGORY_REPORT

@Babi-B, Thanks for the clarity, i appreciate your reviews here.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Babi-B, as per #185 looks like you've not made additions to sources.md. A moment ago i tried to read through and no links or references. I think you should update that. Or would you like me to go ahead with that 🚀🚀🚀🚀🚀

@TimidRobot

Copy link
Copy Markdown
Member

Both use the numbering in the filename to determine order

@Goziee-git The difference is that gcs_fetch.py uses a number in the constant so that they follow the flow when ordered:

FILE1_COUNT
^
FILE2_LANGUAGE
^

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work, thank you!!

@TimidRobot
TimidRobot merged commit 0d44547 into creativecommons:mainNov 1, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Archived in project

Development

Successfully merging this pull request may close these issues.

Integrate arXiv as data source for academic commons quantification

4 participants

@Goziee-git@TimidRobot@Babi-B@cc-open-source-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add arXiv data fetching functionality by Goziee-git · Pull Request #179 · creativecommons/quantifying · GitHub
Skip to content

Add arXiv data fetching functionality - #179

Merged
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv
Nov 1, 2025
Merged

Add arXiv data fetching functionality#179
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv

Conversation

@Goziee-git

@Goziee-gitGoziee-git commented Oct 11, 2025

Copy link
Copy Markdown
Contributor

Fixes

Description

Implements comprehensive arXiv data collection system to quantify open access academic papers in the commons.

Type of Change

  • New feature implementing data collection from the arXivopen access academic papers
  • Data source addition/modification

Changes Made

  • Added arXiv API integration for fetching academic paper metadata using the arXiv API as project requirement for automation of fetching new data sources.
  • Implemented data processing pipeline for arXiv submissions in the scripts/1-fetch/arxiv_fetch.py
  • Created filtering logic for open access and CC-licensed papers
  • Added arXiv data to quarterly reporting system

Testing

  • Static analysis passes (./dev/check.sh)
  • arXiv API integration tested with sample queries
  • Data processing validated with test dataset

Data Impact

  • New data source added (arXiv academic papers)
  • Report generation affected (new academic commons metrics)

Related Documentation

  • Updated sources.md with arXiv API credentials setup
  • Added arXiv processing documentation

Checklist

  • I have read and understood the Developer Certificate of Origin (DCO), below, which covers the contents of this pull request (PR).
  • My pull request doesn't include code or content generated with AI.
  • My pull request has a descriptive title (not a vague title like Update index.md).
  • My pull request targets the default branch of the repository (main or master).
  • My commit messages follow best practices.
  • My code follows the established code style of the repository.
  • I added or updated tests for the changes I made (if applicable).
  • I added or updated documentation (if applicable).
  • I tried running the project locally and verified that there are no
    visible errors.

Developer Certificate of Origin

For the purposes of this DCO, "license" is equivalent to "license or public domain dedication," and "open source license" is equivalent to "open content license or public domain dedication."

Developer Certificate of Origin
Developer Certificate of Origin
Version 1.1
Copyright (C) 2004, 2006 The Linux Foundation and its contributors.
1 Letterman Drive
Suite D4700
San Francisco, CA, 94129
Everyone is permitted to copy and distribute verbatim copies of this
license document, but changing it is not allowed.
Developer's Certificate of Origin 1.1
By making a contribution to this project, I certify that:
(a) The contribution was created in whole or in part by me and I
have the right to submit it under the open source license
indicated in the file; or
(b) The contribution is based upon previous work that, to the best
of my knowledge, is covered under an appropriate open source
license and I have the right under that license to submit that
work with modifications, whether created in whole or in part
by me, under the same open source license (unless I am
permitted to submit under a different license), as indicated
in the file; or
(c) The contribution was provided directly to me by some other
person who certified (a), (b) or (c) and I have not modified
it.
(d) I understand and agree that this project and the contribution
are public and that a record of the contribution (including all
personal information I submit with it, including my sign-off) is
maintained indefinitely and may be redistributed consistent with
this project or the open source license(s) involved.

@Goziee-git
Goziee-git requested review from a team as code ownersOctober 11, 2025 23:29
@Goziee-git
Goziee-git requested review from TimidRobot and possumbilities and removed request for a teamOctober 11, 2025 23:29
@cc-open-source-botcc-open-source-bot moved this to In review in TimidRobotOct 11, 2025
@Goziee-gitGoziee-git changed the title Add arXiv data fetching and processing functionalityAdd arXiv data fetching functionalityOct 12, 2025

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

Comment threadscripts/3-report/gcs_report.py
Comment threaddata/2025Q4/1-fetch/arxiv_1_count.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_category.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_language.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_country.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_year.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_4_count_by_author_count.csv Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

This comment was marked as outdated.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please follow through on your first pull request (PR) before submitting any more:

Depending on how that one goes, I might reopen this PR.

hello @TimidRobot as requested, I have made changes based on your review. Good work is emphasized over speed and I do hope my attempt to go full circle with other PR hasn't dented my chances of significantly contributing to the project. Thank you🙏🏼

@TimidRobotTimidRobot reopened this Oct 17, 2025
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git ok, please focus on this PR

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

@TimidRobot I have removed the query for languages as it returns only English. Also worthy of note here, the arxiv data source accepts papers in other languages but requires that the paper abstracts be submitted in English. so Impossible to get a good distribution of licenses as per language usage.

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time. The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+). The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

Comment threadscripts/3-report/gcs_report.py
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py
Comment threadarxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time.

  • How was arxiv_category_map.yml created?
    • If by script, it should probably go in dev/
  • Data that persists should go in data/ not a specific quarter directory

The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+).

I'll look at data after outstanding comments are resolved.

The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

I'll look at data after outstanding comments are resolved. That said, I'm not excited about adding JSON to the project.

@TimidRobot

Copy link
Copy Markdown
Member

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

It's probably a good idea to create a function in the shared library eventually. Please leave that to last, however.

Refactor arxiv_fetch.py to use requests library for HTTP requests, implementing retry logic for better error handling. Update license extraction logic and CSV headers to remove PLAN_INDEX.
@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot Hello, are there any more changes you'll like to be made to the ones already identified and done. Also i'ld like to ask, if there are no more changes to be made on this, what steps would you recommend as the next one for me in the project, with respect to integrating ArXiv as a valid source for the project?

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Please update sources.md for arXiv

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git Please update sources.md for arXiv

@TimidRobot, Thanks, Im happy we've gotten to this point. very excited to continue contributing to the project. I would like your permission to continue working on the processing and reporting scripts for the arXiv data source

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot as per #179 (comment)
the implementation of the bucketing for the AUTHOR_COUNT was because you suggested that it was not so insightful to just merely dump the data, and so i thought a sort of grouping of this data would help us to generate meaningful insights. Raw author counts would create extremely sparse data with many single-occurrence values while Bucketing ("1", "2-3", "4-6", "7-10", "11+") on the other hand, creates meaningful statistical groups for analysis. It also facilitates visualization and reporting by creating compact, aggregated datasets that are easier to process and analyze hence supporting the project's goal of understanding "how knowledge and culture of the commons is distributed".

although these are my thoughts, I am still very happy to know what you think about this. To further share my opinion, these are some helpful links that guided my decision for bucketing the AUTHOR_COUNT data.

Scientometrics Journal Guidelines: (https://link.springer.com/journal/11192)
• Standard practice in bibliometric studies to group author counts into ranges for collaboration analysis which reduces noise and enables meaningful statistical comparisons across disciplines

FAIR Data Principles: (https://www.go-fair.org/fair-principles/)
• Bucketing supports Findability and Reusability by creating standardized categorical data
NIH Data Management Guidelines: (https://sharing.nih.gov/data-management-and-sharing-policy)

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please see two unresolved conversations:

  1. Add arXiv data fetching functionality #179 (comment)
  2. Add arXiv data fetching functionality #179 (comment)

@TimidRobot the conversations mentioned here are identical, please i'ld like to know what your preference is with the AUTHOR_COUNT, especially if you want it removed or modified in differnt way

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Sorry, let me clarify. I think bucketing authors into AUTHOR_BUCKET is a good and helpful course of action. I wouldd also add https://en.wikipedia.org/wiki/Data_binning to the list of references.

I think which buckets are selected could use some refinement. My thinking is that there are three considerations:

  1. plotting
  2. precedence
  3. distribution

Plotting

For plotting, I think around five values display well.

Precedence

For precedence, we can look at the various citation styles:

Distribution

The script can be modified to provide information on all relevant author counts (only results above 1% are shown):

AuthorsCountPercent
313122.55%
211119.10%
49215.83%
18514.63%
5467.92%
6233.96%
7213.61%
8193.27%
9152.58%
10101.72%

Recommendation

I recommend the following buckets:

  • 1 author
  • 2 authors
  • 3 authors
  • 4 authors
  • 5+ authors

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@Goziee-git

Goziee-git commented Oct 31, 2025

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot as per ordering the constants, what would be your preference here. The current order of the constants is:

  1. FILE_ARXIV_COUNT (arxiv_1_count.csv)
  2. FILE_ARXIV_CATEGORY_REPORT (arxiv_2_count_by_category_report.csv)
  3. FILE_ARXIV_YEAR (arxiv_3_count_by_year.csv)
  4. FILE_ARXIV_AUTHOR_BUCKET (arxiv_4_count_by_author_bucket.csv)

This follows the logical/sequential order based on the numbered workflow (1 → 2 → 3 → 4).

in comparison to the gcs_fetch.py (current order):

  1. FILE1_COUNT (gcs_1_count.csv)
  2. FILE2_LANGUAGE (gcs_2_count_by_language.csv)
  3. FILE3_COUNTRY (gcs_3_count_by_country.csv)

Comparison:

  • Both scripts follow the same logical/sequential ordering pattern
  • Both order constants by their numbered workflow sequence (1 → 2 → 3 → 4)
  • Both use the numbering in the filename to determine order

Do you prefer some other order for this, so i can be guided as to the best course of action here

@Babi-B

Babi-B commented Oct 31, 2025

Copy link
Copy Markdown
Contributor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before
FILE_ARXIV_CATEGORY_REPORT

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before FILE_ARXIV_CATEGORY_REPORT

@Babi-B, Thanks for the clarity, i appreciate your reviews here.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Babi-B, as per #185 looks like you've not made additions to sources.md. A moment ago i tried to read through and no links or references. I think you should update that. Or would you like me to go ahead with that 🚀🚀🚀🚀🚀

@TimidRobot

Copy link
Copy Markdown
Member

Both use the numbering in the filename to determine order

@Goziee-git The difference is that gcs_fetch.py uses a number in the constant so that they follow the flow when ordered:

FILE1_COUNT
^
FILE2_LANGUAGE
^

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work, thank you!!

@TimidRobot
TimidRobot merged commit 0d44547 into creativecommons:mainNov 1, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Archived in project

Development

Successfully merging this pull request may close these issues.

Integrate arXiv as data source for academic commons quantification

4 participants

@Goziee-git@TimidRobot@Babi-B@cc-open-source-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Add arXiv data fetching functionality by Goziee-git · Pull Request #179 · creativecommons/quantifying · GitHub
Skip to content

Add arXiv data fetching functionality - #179

Merged
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv
Nov 1, 2025
Merged

Add arXiv data fetching functionality#179
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv

Conversation

@Goziee-git

@Goziee-gitGoziee-git commented Oct 11, 2025

Copy link
Copy Markdown
Contributor

Fixes

Description

Implements comprehensive arXiv data collection system to quantify open access academic papers in the commons.

Type of Change

  • New feature implementing data collection from the arXivopen access academic papers
  • Data source addition/modification

Changes Made

  • Added arXiv API integration for fetching academic paper metadata using the arXiv API as project requirement for automation of fetching new data sources.
  • Implemented data processing pipeline for arXiv submissions in the scripts/1-fetch/arxiv_fetch.py
  • Created filtering logic for open access and CC-licensed papers
  • Added arXiv data to quarterly reporting system

Testing

  • Static analysis passes (./dev/check.sh)
  • arXiv API integration tested with sample queries
  • Data processing validated with test dataset

Data Impact

  • New data source added (arXiv academic papers)
  • Report generation affected (new academic commons metrics)

Related Documentation

  • Updated sources.md with arXiv API credentials setup
  • Added arXiv processing documentation

Checklist

  • I have read and understood the Developer Certificate of Origin (DCO), below, which covers the contents of this pull request (PR).
  • My pull request doesn't include code or content generated with AI.
  • My pull request has a descriptive title (not a vague title like Update index.md).
  • My pull request targets the default branch of the repository (main or master).
  • My commit messages follow best practices.
  • My code follows the established code style of the repository.
  • I added or updated tests for the changes I made (if applicable).
  • I added or updated documentation (if applicable).
  • I tried running the project locally and verified that there are no
    visible errors.

Developer Certificate of Origin

For the purposes of this DCO, "license" is equivalent to "license or public domain dedication," and "open source license" is equivalent to "open content license or public domain dedication."

Developer Certificate of Origin
Developer Certificate of Origin
Version 1.1
Copyright (C) 2004, 2006 The Linux Foundation and its contributors.
1 Letterman Drive
Suite D4700
San Francisco, CA, 94129
Everyone is permitted to copy and distribute verbatim copies of this
license document, but changing it is not allowed.
Developer's Certificate of Origin 1.1
By making a contribution to this project, I certify that:
(a) The contribution was created in whole or in part by me and I
have the right to submit it under the open source license
indicated in the file; or
(b) The contribution is based upon previous work that, to the best
of my knowledge, is covered under an appropriate open source
license and I have the right under that license to submit that
work with modifications, whether created in whole or in part
by me, under the same open source license (unless I am
permitted to submit under a different license), as indicated
in the file; or
(c) The contribution was provided directly to me by some other
person who certified (a), (b) or (c) and I have not modified
it.
(d) I understand and agree that this project and the contribution
are public and that a record of the contribution (including all
personal information I submit with it, including my sign-off) is
maintained indefinitely and may be redistributed consistent with
this project or the open source license(s) involved.

@Goziee-git
Goziee-git requested review from a team as code ownersOctober 11, 2025 23:29
@Goziee-git
Goziee-git requested review from TimidRobot and possumbilities and removed request for a teamOctober 11, 2025 23:29
@cc-open-source-botcc-open-source-bot moved this to In review in TimidRobotOct 11, 2025
@Goziee-gitGoziee-git changed the title Add arXiv data fetching and processing functionalityAdd arXiv data fetching functionalityOct 12, 2025

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

Comment threadscripts/3-report/gcs_report.py
Comment threaddata/2025Q4/1-fetch/arxiv_1_count.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_category.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_language.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_country.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_year.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_4_count_by_author_count.csv Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

This comment was marked as outdated.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please follow through on your first pull request (PR) before submitting any more:

Depending on how that one goes, I might reopen this PR.

hello @TimidRobot as requested, I have made changes based on your review. Good work is emphasized over speed and I do hope my attempt to go full circle with other PR hasn't dented my chances of significantly contributing to the project. Thank you🙏🏼

@TimidRobotTimidRobot reopened this Oct 17, 2025
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git ok, please focus on this PR

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

@TimidRobot I have removed the query for languages as it returns only English. Also worthy of note here, the arxiv data source accepts papers in other languages but requires that the paper abstracts be submitted in English. so Impossible to get a good distribution of licenses as per language usage.

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time. The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+). The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

Comment threadscripts/3-report/gcs_report.py
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py
Comment threadarxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time.

  • How was arxiv_category_map.yml created?
    • If by script, it should probably go in dev/
  • Data that persists should go in data/ not a specific quarter directory

The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+).

I'll look at data after outstanding comments are resolved.

The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

I'll look at data after outstanding comments are resolved. That said, I'm not excited about adding JSON to the project.

@TimidRobot

Copy link
Copy Markdown
Member

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

It's probably a good idea to create a function in the shared library eventually. Please leave that to last, however.

Refactor arxiv_fetch.py to use requests library for HTTP requests, implementing retry logic for better error handling. Update license extraction logic and CSV headers to remove PLAN_INDEX.
@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot Hello, are there any more changes you'll like to be made to the ones already identified and done. Also i'ld like to ask, if there are no more changes to be made on this, what steps would you recommend as the next one for me in the project, with respect to integrating ArXiv as a valid source for the project?

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Please update sources.md for arXiv

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git Please update sources.md for arXiv

@TimidRobot, Thanks, Im happy we've gotten to this point. very excited to continue contributing to the project. I would like your permission to continue working on the processing and reporting scripts for the arXiv data source

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot as per #179 (comment)
the implementation of the bucketing for the AUTHOR_COUNT was because you suggested that it was not so insightful to just merely dump the data, and so i thought a sort of grouping of this data would help us to generate meaningful insights. Raw author counts would create extremely sparse data with many single-occurrence values while Bucketing ("1", "2-3", "4-6", "7-10", "11+") on the other hand, creates meaningful statistical groups for analysis. It also facilitates visualization and reporting by creating compact, aggregated datasets that are easier to process and analyze hence supporting the project's goal of understanding "how knowledge and culture of the commons is distributed".

although these are my thoughts, I am still very happy to know what you think about this. To further share my opinion, these are some helpful links that guided my decision for bucketing the AUTHOR_COUNT data.

Scientometrics Journal Guidelines: (https://link.springer.com/journal/11192)
• Standard practice in bibliometric studies to group author counts into ranges for collaboration analysis which reduces noise and enables meaningful statistical comparisons across disciplines

FAIR Data Principles: (https://www.go-fair.org/fair-principles/)
• Bucketing supports Findability and Reusability by creating standardized categorical data
NIH Data Management Guidelines: (https://sharing.nih.gov/data-management-and-sharing-policy)

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please see two unresolved conversations:

  1. Add arXiv data fetching functionality #179 (comment)
  2. Add arXiv data fetching functionality #179 (comment)

@TimidRobot the conversations mentioned here are identical, please i'ld like to know what your preference is with the AUTHOR_COUNT, especially if you want it removed or modified in differnt way

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Sorry, let me clarify. I think bucketing authors into AUTHOR_BUCKET is a good and helpful course of action. I wouldd also add https://en.wikipedia.org/wiki/Data_binning to the list of references.

I think which buckets are selected could use some refinement. My thinking is that there are three considerations:

  1. plotting
  2. precedence
  3. distribution

Plotting

For plotting, I think around five values display well.

Precedence

For precedence, we can look at the various citation styles:

Distribution

The script can be modified to provide information on all relevant author counts (only results above 1% are shown):

AuthorsCountPercent
313122.55%
211119.10%
49215.83%
18514.63%
5467.92%
6233.96%
7213.61%
8193.27%
9152.58%
10101.72%

Recommendation

I recommend the following buckets:

  • 1 author
  • 2 authors
  • 3 authors
  • 4 authors
  • 5+ authors

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@Goziee-git

Goziee-git commented Oct 31, 2025

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot as per ordering the constants, what would be your preference here. The current order of the constants is:

  1. FILE_ARXIV_COUNT (arxiv_1_count.csv)
  2. FILE_ARXIV_CATEGORY_REPORT (arxiv_2_count_by_category_report.csv)
  3. FILE_ARXIV_YEAR (arxiv_3_count_by_year.csv)
  4. FILE_ARXIV_AUTHOR_BUCKET (arxiv_4_count_by_author_bucket.csv)

This follows the logical/sequential order based on the numbered workflow (1 → 2 → 3 → 4).

in comparison to the gcs_fetch.py (current order):

  1. FILE1_COUNT (gcs_1_count.csv)
  2. FILE2_LANGUAGE (gcs_2_count_by_language.csv)
  3. FILE3_COUNTRY (gcs_3_count_by_country.csv)

Comparison:

  • Both scripts follow the same logical/sequential ordering pattern
  • Both order constants by their numbered workflow sequence (1 → 2 → 3 → 4)
  • Both use the numbering in the filename to determine order

Do you prefer some other order for this, so i can be guided as to the best course of action here

@Babi-B

Babi-B commented Oct 31, 2025

Copy link
Copy Markdown
Contributor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before
FILE_ARXIV_CATEGORY_REPORT

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before FILE_ARXIV_CATEGORY_REPORT

@Babi-B, Thanks for the clarity, i appreciate your reviews here.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Babi-B, as per #185 looks like you've not made additions to sources.md. A moment ago i tried to read through and no links or references. I think you should update that. Or would you like me to go ahead with that 🚀🚀🚀🚀🚀

@TimidRobot

Copy link
Copy Markdown
Member

Both use the numbering in the filename to determine order

@Goziee-git The difference is that gcs_fetch.py uses a number in the constant so that they follow the flow when ordered:

FILE1_COUNT
^
FILE2_LANGUAGE
^

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work, thank you!!

@TimidRobot
TimidRobot merged commit 0d44547 into creativecommons:mainNov 1, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Archived in project

Development

Successfully merging this pull request may close these issues.

Integrate arXiv as data source for academic commons quantification

4 participants

@Goziee-git@TimidRobot@Babi-B@cc-open-source-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add arXiv data fetching functionality by Goziee-git · Pull Request #179 · creativecommons/quantifying · GitHub
Skip to content

Add arXiv data fetching functionality - #179

Merged
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv
Nov 1, 2025
Merged

Add arXiv data fetching functionality#179
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv

Conversation

@Goziee-git

@Goziee-gitGoziee-git commented Oct 11, 2025

Copy link
Copy Markdown
Contributor

Fixes

Description

Implements comprehensive arXiv data collection system to quantify open access academic papers in the commons.

Type of Change

  • New feature implementing data collection from the arXivopen access academic papers
  • Data source addition/modification

Changes Made

  • Added arXiv API integration for fetching academic paper metadata using the arXiv API as project requirement for automation of fetching new data sources.
  • Implemented data processing pipeline for arXiv submissions in the scripts/1-fetch/arxiv_fetch.py
  • Created filtering logic for open access and CC-licensed papers
  • Added arXiv data to quarterly reporting system

Testing

  • Static analysis passes (./dev/check.sh)
  • arXiv API integration tested with sample queries
  • Data processing validated with test dataset

Data Impact

  • New data source added (arXiv academic papers)
  • Report generation affected (new academic commons metrics)

Related Documentation

  • Updated sources.md with arXiv API credentials setup
  • Added arXiv processing documentation

Checklist

  • I have read and understood the Developer Certificate of Origin (DCO), below, which covers the contents of this pull request (PR).
  • My pull request doesn't include code or content generated with AI.
  • My pull request has a descriptive title (not a vague title like Update index.md).
  • My pull request targets the default branch of the repository (main or master).
  • My commit messages follow best practices.
  • My code follows the established code style of the repository.
  • I added or updated tests for the changes I made (if applicable).
  • I added or updated documentation (if applicable).
  • I tried running the project locally and verified that there are no
    visible errors.

Developer Certificate of Origin

For the purposes of this DCO, "license" is equivalent to "license or public domain dedication," and "open source license" is equivalent to "open content license or public domain dedication."

Developer Certificate of Origin
Developer Certificate of Origin
Version 1.1
Copyright (C) 2004, 2006 The Linux Foundation and its contributors.
1 Letterman Drive
Suite D4700
San Francisco, CA, 94129
Everyone is permitted to copy and distribute verbatim copies of this
license document, but changing it is not allowed.
Developer's Certificate of Origin 1.1
By making a contribution to this project, I certify that:
(a) The contribution was created in whole or in part by me and I
have the right to submit it under the open source license
indicated in the file; or
(b) The contribution is based upon previous work that, to the best
of my knowledge, is covered under an appropriate open source
license and I have the right under that license to submit that
work with modifications, whether created in whole or in part
by me, under the same open source license (unless I am
permitted to submit under a different license), as indicated
in the file; or
(c) The contribution was provided directly to me by some other
person who certified (a), (b) or (c) and I have not modified
it.
(d) I understand and agree that this project and the contribution
are public and that a record of the contribution (including all
personal information I submit with it, including my sign-off) is
maintained indefinitely and may be redistributed consistent with
this project or the open source license(s) involved.

@Goziee-git
Goziee-git requested review from a team as code ownersOctober 11, 2025 23:29
@Goziee-git
Goziee-git requested review from TimidRobot and possumbilities and removed request for a teamOctober 11, 2025 23:29
@cc-open-source-botcc-open-source-bot moved this to In review in TimidRobotOct 11, 2025
@Goziee-gitGoziee-git changed the title Add arXiv data fetching and processing functionalityAdd arXiv data fetching functionalityOct 12, 2025

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

Comment threadscripts/3-report/gcs_report.py
Comment threaddata/2025Q4/1-fetch/arxiv_1_count.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_category.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_language.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_country.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_year.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_4_count_by_author_count.csv Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

This comment was marked as outdated.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please follow through on your first pull request (PR) before submitting any more:

Depending on how that one goes, I might reopen this PR.

hello @TimidRobot as requested, I have made changes based on your review. Good work is emphasized over speed and I do hope my attempt to go full circle with other PR hasn't dented my chances of significantly contributing to the project. Thank you🙏🏼

@TimidRobotTimidRobot reopened this Oct 17, 2025
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git ok, please focus on this PR

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

@TimidRobot I have removed the query for languages as it returns only English. Also worthy of note here, the arxiv data source accepts papers in other languages but requires that the paper abstracts be submitted in English. so Impossible to get a good distribution of licenses as per language usage.

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time. The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+). The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

Comment threadscripts/3-report/gcs_report.py
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py
Comment threadarxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time.

  • How was arxiv_category_map.yml created?
    • If by script, it should probably go in dev/
  • Data that persists should go in data/ not a specific quarter directory

The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+).

I'll look at data after outstanding comments are resolved.

The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

I'll look at data after outstanding comments are resolved. That said, I'm not excited about adding JSON to the project.

@TimidRobot

Copy link
Copy Markdown
Member

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

It's probably a good idea to create a function in the shared library eventually. Please leave that to last, however.

Refactor arxiv_fetch.py to use requests library for HTTP requests, implementing retry logic for better error handling. Update license extraction logic and CSV headers to remove PLAN_INDEX.
@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot Hello, are there any more changes you'll like to be made to the ones already identified and done. Also i'ld like to ask, if there are no more changes to be made on this, what steps would you recommend as the next one for me in the project, with respect to integrating ArXiv as a valid source for the project?

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Please update sources.md for arXiv

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git Please update sources.md for arXiv

@TimidRobot, Thanks, Im happy we've gotten to this point. very excited to continue contributing to the project. I would like your permission to continue working on the processing and reporting scripts for the arXiv data source

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot as per #179 (comment)
the implementation of the bucketing for the AUTHOR_COUNT was because you suggested that it was not so insightful to just merely dump the data, and so i thought a sort of grouping of this data would help us to generate meaningful insights. Raw author counts would create extremely sparse data with many single-occurrence values while Bucketing ("1", "2-3", "4-6", "7-10", "11+") on the other hand, creates meaningful statistical groups for analysis. It also facilitates visualization and reporting by creating compact, aggregated datasets that are easier to process and analyze hence supporting the project's goal of understanding "how knowledge and culture of the commons is distributed".

although these are my thoughts, I am still very happy to know what you think about this. To further share my opinion, these are some helpful links that guided my decision for bucketing the AUTHOR_COUNT data.

Scientometrics Journal Guidelines: (https://link.springer.com/journal/11192)
• Standard practice in bibliometric studies to group author counts into ranges for collaboration analysis which reduces noise and enables meaningful statistical comparisons across disciplines

FAIR Data Principles: (https://www.go-fair.org/fair-principles/)
• Bucketing supports Findability and Reusability by creating standardized categorical data
NIH Data Management Guidelines: (https://sharing.nih.gov/data-management-and-sharing-policy)

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please see two unresolved conversations:

  1. Add arXiv data fetching functionality #179 (comment)
  2. Add arXiv data fetching functionality #179 (comment)

@TimidRobot the conversations mentioned here are identical, please i'ld like to know what your preference is with the AUTHOR_COUNT, especially if you want it removed or modified in differnt way

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Sorry, let me clarify. I think bucketing authors into AUTHOR_BUCKET is a good and helpful course of action. I wouldd also add https://en.wikipedia.org/wiki/Data_binning to the list of references.

I think which buckets are selected could use some refinement. My thinking is that there are three considerations:

  1. plotting
  2. precedence
  3. distribution

Plotting

For plotting, I think around five values display well.

Precedence

For precedence, we can look at the various citation styles:

Distribution

The script can be modified to provide information on all relevant author counts (only results above 1% are shown):

AuthorsCountPercent
313122.55%
211119.10%
49215.83%
18514.63%
5467.92%
6233.96%
7213.61%
8193.27%
9152.58%
10101.72%

Recommendation

I recommend the following buckets:

  • 1 author
  • 2 authors
  • 3 authors
  • 4 authors
  • 5+ authors

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@Goziee-git

Goziee-git commented Oct 31, 2025

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot as per ordering the constants, what would be your preference here. The current order of the constants is:

  1. FILE_ARXIV_COUNT (arxiv_1_count.csv)
  2. FILE_ARXIV_CATEGORY_REPORT (arxiv_2_count_by_category_report.csv)
  3. FILE_ARXIV_YEAR (arxiv_3_count_by_year.csv)
  4. FILE_ARXIV_AUTHOR_BUCKET (arxiv_4_count_by_author_bucket.csv)

This follows the logical/sequential order based on the numbered workflow (1 → 2 → 3 → 4).

in comparison to the gcs_fetch.py (current order):

  1. FILE1_COUNT (gcs_1_count.csv)
  2. FILE2_LANGUAGE (gcs_2_count_by_language.csv)
  3. FILE3_COUNTRY (gcs_3_count_by_country.csv)

Comparison:

  • Both scripts follow the same logical/sequential ordering pattern
  • Both order constants by their numbered workflow sequence (1 → 2 → 3 → 4)
  • Both use the numbering in the filename to determine order

Do you prefer some other order for this, so i can be guided as to the best course of action here

@Babi-B

Babi-B commented Oct 31, 2025

Copy link
Copy Markdown
Contributor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before
FILE_ARXIV_CATEGORY_REPORT

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before FILE_ARXIV_CATEGORY_REPORT

@Babi-B, Thanks for the clarity, i appreciate your reviews here.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Babi-B, as per #185 looks like you've not made additions to sources.md. A moment ago i tried to read through and no links or references. I think you should update that. Or would you like me to go ahead with that 🚀🚀🚀🚀🚀

@TimidRobot

Copy link
Copy Markdown
Member

Both use the numbering in the filename to determine order

@Goziee-git The difference is that gcs_fetch.py uses a number in the constant so that they follow the flow when ordered:

FILE1_COUNT
^
FILE2_LANGUAGE
^

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work, thank you!!

@TimidRobot
TimidRobot merged commit 0d44547 into creativecommons:mainNov 1, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Archived in project

Development

Successfully merging this pull request may close these issues.

Integrate arXiv as data source for academic commons quantification

4 participants

@Goziee-git@TimidRobot@Babi-B@cc-open-source-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add arXiv data fetching functionality by Goziee-git · Pull Request #179 · creativecommons/quantifying · GitHub
Skip to content

Add arXiv data fetching functionality - #179

Merged
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv
Nov 1, 2025
Merged

Add arXiv data fetching functionality#179
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv

Conversation

@Goziee-git

@Goziee-gitGoziee-git commented Oct 11, 2025

Copy link
Copy Markdown
Contributor

Fixes

Description

Implements comprehensive arXiv data collection system to quantify open access academic papers in the commons.

Type of Change

  • New feature implementing data collection from the arXivopen access academic papers
  • Data source addition/modification

Changes Made

  • Added arXiv API integration for fetching academic paper metadata using the arXiv API as project requirement for automation of fetching new data sources.
  • Implemented data processing pipeline for arXiv submissions in the scripts/1-fetch/arxiv_fetch.py
  • Created filtering logic for open access and CC-licensed papers
  • Added arXiv data to quarterly reporting system

Testing

  • Static analysis passes (./dev/check.sh)
  • arXiv API integration tested with sample queries
  • Data processing validated with test dataset

Data Impact

  • New data source added (arXiv academic papers)
  • Report generation affected (new academic commons metrics)

Related Documentation

  • Updated sources.md with arXiv API credentials setup
  • Added arXiv processing documentation

Checklist

  • I have read and understood the Developer Certificate of Origin (DCO), below, which covers the contents of this pull request (PR).
  • My pull request doesn't include code or content generated with AI.
  • My pull request has a descriptive title (not a vague title like Update index.md).
  • My pull request targets the default branch of the repository (main or master).
  • My commit messages follow best practices.
  • My code follows the established code style of the repository.
  • I added or updated tests for the changes I made (if applicable).
  • I added or updated documentation (if applicable).
  • I tried running the project locally and verified that there are no
    visible errors.

Developer Certificate of Origin

For the purposes of this DCO, "license" is equivalent to "license or public domain dedication," and "open source license" is equivalent to "open content license or public domain dedication."

Developer Certificate of Origin
Developer Certificate of Origin
Version 1.1
Copyright (C) 2004, 2006 The Linux Foundation and its contributors.
1 Letterman Drive
Suite D4700
San Francisco, CA, 94129
Everyone is permitted to copy and distribute verbatim copies of this
license document, but changing it is not allowed.
Developer's Certificate of Origin 1.1
By making a contribution to this project, I certify that:
(a) The contribution was created in whole or in part by me and I
have the right to submit it under the open source license
indicated in the file; or
(b) The contribution is based upon previous work that, to the best
of my knowledge, is covered under an appropriate open source
license and I have the right under that license to submit that
work with modifications, whether created in whole or in part
by me, under the same open source license (unless I am
permitted to submit under a different license), as indicated
in the file; or
(c) The contribution was provided directly to me by some other
person who certified (a), (b) or (c) and I have not modified
it.
(d) I understand and agree that this project and the contribution
are public and that a record of the contribution (including all
personal information I submit with it, including my sign-off) is
maintained indefinitely and may be redistributed consistent with
this project or the open source license(s) involved.

@Goziee-git
Goziee-git requested review from a team as code ownersOctober 11, 2025 23:29
@Goziee-git
Goziee-git requested review from TimidRobot and possumbilities and removed request for a teamOctober 11, 2025 23:29
@cc-open-source-botcc-open-source-bot moved this to In review in TimidRobotOct 11, 2025
@Goziee-gitGoziee-git changed the title Add arXiv data fetching and processing functionalityAdd arXiv data fetching functionalityOct 12, 2025

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

Comment threadscripts/3-report/gcs_report.py
Comment threaddata/2025Q4/1-fetch/arxiv_1_count.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_category.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_language.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_country.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_year.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_4_count_by_author_count.csv Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

This comment was marked as outdated.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please follow through on your first pull request (PR) before submitting any more:

Depending on how that one goes, I might reopen this PR.

hello @TimidRobot as requested, I have made changes based on your review. Good work is emphasized over speed and I do hope my attempt to go full circle with other PR hasn't dented my chances of significantly contributing to the project. Thank you🙏🏼

@TimidRobotTimidRobot reopened this Oct 17, 2025
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git ok, please focus on this PR

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

@TimidRobot I have removed the query for languages as it returns only English. Also worthy of note here, the arxiv data source accepts papers in other languages but requires that the paper abstracts be submitted in English. so Impossible to get a good distribution of licenses as per language usage.

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time. The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+). The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

Comment threadscripts/3-report/gcs_report.py
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py
Comment threadarxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time.

  • How was arxiv_category_map.yml created?
    • If by script, it should probably go in dev/
  • Data that persists should go in data/ not a specific quarter directory

The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+).

I'll look at data after outstanding comments are resolved.

The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

I'll look at data after outstanding comments are resolved. That said, I'm not excited about adding JSON to the project.

@TimidRobot

Copy link
Copy Markdown
Member

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

It's probably a good idea to create a function in the shared library eventually. Please leave that to last, however.

Refactor arxiv_fetch.py to use requests library for HTTP requests, implementing retry logic for better error handling. Update license extraction logic and CSV headers to remove PLAN_INDEX.
@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot Hello, are there any more changes you'll like to be made to the ones already identified and done. Also i'ld like to ask, if there are no more changes to be made on this, what steps would you recommend as the next one for me in the project, with respect to integrating ArXiv as a valid source for the project?

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Please update sources.md for arXiv

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git Please update sources.md for arXiv

@TimidRobot, Thanks, Im happy we've gotten to this point. very excited to continue contributing to the project. I would like your permission to continue working on the processing and reporting scripts for the arXiv data source

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot as per #179 (comment)
the implementation of the bucketing for the AUTHOR_COUNT was because you suggested that it was not so insightful to just merely dump the data, and so i thought a sort of grouping of this data would help us to generate meaningful insights. Raw author counts would create extremely sparse data with many single-occurrence values while Bucketing ("1", "2-3", "4-6", "7-10", "11+") on the other hand, creates meaningful statistical groups for analysis. It also facilitates visualization and reporting by creating compact, aggregated datasets that are easier to process and analyze hence supporting the project's goal of understanding "how knowledge and culture of the commons is distributed".

although these are my thoughts, I am still very happy to know what you think about this. To further share my opinion, these are some helpful links that guided my decision for bucketing the AUTHOR_COUNT data.

Scientometrics Journal Guidelines: (https://link.springer.com/journal/11192)
• Standard practice in bibliometric studies to group author counts into ranges for collaboration analysis which reduces noise and enables meaningful statistical comparisons across disciplines

FAIR Data Principles: (https://www.go-fair.org/fair-principles/)
• Bucketing supports Findability and Reusability by creating standardized categorical data
NIH Data Management Guidelines: (https://sharing.nih.gov/data-management-and-sharing-policy)

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please see two unresolved conversations:

  1. Add arXiv data fetching functionality #179 (comment)
  2. Add arXiv data fetching functionality #179 (comment)

@TimidRobot the conversations mentioned here are identical, please i'ld like to know what your preference is with the AUTHOR_COUNT, especially if you want it removed or modified in differnt way

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Sorry, let me clarify. I think bucketing authors into AUTHOR_BUCKET is a good and helpful course of action. I wouldd also add https://en.wikipedia.org/wiki/Data_binning to the list of references.

I think which buckets are selected could use some refinement. My thinking is that there are three considerations:

  1. plotting
  2. precedence
  3. distribution

Plotting

For plotting, I think around five values display well.

Precedence

For precedence, we can look at the various citation styles:

Distribution

The script can be modified to provide information on all relevant author counts (only results above 1% are shown):

AuthorsCountPercent
313122.55%
211119.10%
49215.83%
18514.63%
5467.92%
6233.96%
7213.61%
8193.27%
9152.58%
10101.72%

Recommendation

I recommend the following buckets:

  • 1 author
  • 2 authors
  • 3 authors
  • 4 authors
  • 5+ authors

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@Goziee-git

Goziee-git commented Oct 31, 2025

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot as per ordering the constants, what would be your preference here. The current order of the constants is:

  1. FILE_ARXIV_COUNT (arxiv_1_count.csv)
  2. FILE_ARXIV_CATEGORY_REPORT (arxiv_2_count_by_category_report.csv)
  3. FILE_ARXIV_YEAR (arxiv_3_count_by_year.csv)
  4. FILE_ARXIV_AUTHOR_BUCKET (arxiv_4_count_by_author_bucket.csv)

This follows the logical/sequential order based on the numbered workflow (1 → 2 → 3 → 4).

in comparison to the gcs_fetch.py (current order):

  1. FILE1_COUNT (gcs_1_count.csv)
  2. FILE2_LANGUAGE (gcs_2_count_by_language.csv)
  3. FILE3_COUNTRY (gcs_3_count_by_country.csv)

Comparison:

  • Both scripts follow the same logical/sequential ordering pattern
  • Both order constants by their numbered workflow sequence (1 → 2 → 3 → 4)
  • Both use the numbering in the filename to determine order

Do you prefer some other order for this, so i can be guided as to the best course of action here

@Babi-B

Babi-B commented Oct 31, 2025

Copy link
Copy Markdown
Contributor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before
FILE_ARXIV_CATEGORY_REPORT

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before FILE_ARXIV_CATEGORY_REPORT

@Babi-B, Thanks for the clarity, i appreciate your reviews here.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Babi-B, as per #185 looks like you've not made additions to sources.md. A moment ago i tried to read through and no links or references. I think you should update that. Or would you like me to go ahead with that 🚀🚀🚀🚀🚀

@TimidRobot

Copy link
Copy Markdown
Member

Both use the numbering in the filename to determine order

@Goziee-git The difference is that gcs_fetch.py uses a number in the constant so that they follow the flow when ordered:

FILE1_COUNT
^
FILE2_LANGUAGE
^

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work, thank you!!

@TimidRobot
TimidRobot merged commit 0d44547 into creativecommons:mainNov 1, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Archived in project

Development

Successfully merging this pull request may close these issues.

Integrate arXiv as data source for academic commons quantification

4 participants

@Goziee-git@TimidRobot@Babi-B@cc-open-source-bot
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Add arXiv data fetching functionality by Goziee-git · Pull Request #179 · creativecommons/quantifying · GitHub
Skip to content

Add arXiv data fetching functionality - #179

Merged
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv
Nov 1, 2025
Merged

Add arXiv data fetching functionality#179
TimidRobot merged 64 commits into
creativecommons:mainfrom
Goziee-git:feature/arxiv

Conversation

@Goziee-git

@Goziee-gitGoziee-git commented Oct 11, 2025

Copy link
Copy Markdown
Contributor

Fixes

Description

Implements comprehensive arXiv data collection system to quantify open access academic papers in the commons.

Type of Change

  • New feature implementing data collection from the arXivopen access academic papers
  • Data source addition/modification

Changes Made

  • Added arXiv API integration for fetching academic paper metadata using the arXiv API as project requirement for automation of fetching new data sources.
  • Implemented data processing pipeline for arXiv submissions in the scripts/1-fetch/arxiv_fetch.py
  • Created filtering logic for open access and CC-licensed papers
  • Added arXiv data to quarterly reporting system

Testing

  • Static analysis passes (./dev/check.sh)
  • arXiv API integration tested with sample queries
  • Data processing validated with test dataset

Data Impact

  • New data source added (arXiv academic papers)
  • Report generation affected (new academic commons metrics)

Related Documentation

  • Updated sources.md with arXiv API credentials setup
  • Added arXiv processing documentation

Checklist

  • I have read and understood the Developer Certificate of Origin (DCO), below, which covers the contents of this pull request (PR).
  • My pull request doesn't include code or content generated with AI.
  • My pull request has a descriptive title (not a vague title like Update index.md).
  • My pull request targets the default branch of the repository (main or master).
  • My commit messages follow best practices.
  • My code follows the established code style of the repository.
  • I added or updated tests for the changes I made (if applicable).
  • I added or updated documentation (if applicable).
  • I tried running the project locally and verified that there are no
    visible errors.

Developer Certificate of Origin

For the purposes of this DCO, "license" is equivalent to "license or public domain dedication," and "open source license" is equivalent to "open content license or public domain dedication."

Developer Certificate of Origin
Developer Certificate of Origin
Version 1.1
Copyright (C) 2004, 2006 The Linux Foundation and its contributors.
1 Letterman Drive
Suite D4700
San Francisco, CA, 94129
Everyone is permitted to copy and distribute verbatim copies of this
license document, but changing it is not allowed.
Developer's Certificate of Origin 1.1
By making a contribution to this project, I certify that:
(a) The contribution was created in whole or in part by me and I
have the right to submit it under the open source license
indicated in the file; or
(b) The contribution is based upon previous work that, to the best
of my knowledge, is covered under an appropriate open source
license and I have the right under that license to submit that
work with modifications, whether created in whole or in part
by me, under the same open source license (unless I am
permitted to submit under a different license), as indicated
in the file; or
(c) The contribution was provided directly to me by some other
person who certified (a), (b) or (c) and I have not modified
it.
(d) I understand and agree that this project and the contribution
are public and that a record of the contribution (including all
personal information I submit with it, including my sign-off) is
maintained indefinitely and may be redistributed consistent with
this project or the open source license(s) involved.

@Goziee-git
Goziee-git requested review from a team as code ownersOctober 11, 2025 23:29
@Goziee-git
Goziee-git requested review from TimidRobot and possumbilities and removed request for a teamOctober 11, 2025 23:29
@cc-open-source-botcc-open-source-bot moved this to In review in TimidRobotOct 11, 2025
@Goziee-gitGoziee-git changed the title Add arXiv data fetching and processing functionalityAdd arXiv data fetching functionalityOct 12, 2025

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

Comment threadscripts/3-report/gcs_report.py
Comment threaddata/2025Q4/1-fetch/arxiv_1_count.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_category.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_2_count_by_language.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_country.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_3_count_by_year.csv Outdated
Comment threaddata/2025Q4/1-fetch/arxiv_4_count_by_author_count.csv Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

This comment was marked as outdated.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please follow through on your first pull request (PR) before submitting any more:

Depending on how that one goes, I might reopen this PR.

hello @TimidRobot as requested, I have made changes based on your review. Good work is emphasized over speed and I do hope my attempt to go full circle with other PR hasn't dented my chances of significantly contributing to the project. Thank you🙏🏼

@TimidRobotTimidRobot reopened this Oct 17, 2025
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git ok, please focus on this PR

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

This is a great start.

I recommend also developing a data/report plan. For example:

  • It is not meaningful to get a count of a single language (though it is worth noting that other languages are not available).
  • Category codes converted to reporting (words and/or abbreviations instead of acronyms)

@TimidRobot I have removed the query for languages as it returns only English. Also worthy of note here, the arxiv data source accepts papers in other languages but requires that the paper abstracts be submitted in English. so Impossible to get a good distribution of licenses as per language usage.

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time. The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+). The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

@Goziee-git

Goziee-git commented Oct 20, 2025

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

Comment threadscripts/3-report/gcs_report.py
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
Comment threadscripts/1-fetch/arxiv_fetch.py
Comment threadarxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

Also, as suggested. I converted the category codes to reporting words that a more user-friendly and readable using an external arxiv_category_map.yml, in data/2025Q4/1-fetch. I believe this should make updates repoducible and maintainable over time.

  • How was arxiv_category_map.yml created?
    • If by script, it should probably go in dev/
  • Data that persists should go in data/ not a specific quarter directory

The script now produces arxiv_2_count_by_category_report.csv and arxiv_2_count_by_category_report_agg.csv for better reporting. Also instead of dumping author count data previously, i implemented a Bucketing approach. in arxiv_4_count_by_author_bucket.csv to group author counts into meaningful ranges (1, 2-3, 4-6, 7-10, 11+).

I'll look at data after outstanding comments are resolved.

The script also generates a arxiv_provenance.json to record metadata for audit, reproducibility, and provenance

I'll look at data after outstanding comments are resolved. That said, I'm not excited about adding JSON to the project.

@TimidRobot

Copy link
Copy Markdown
Member

Hello @TimidRobot, i observed from multiple results fetched previously that the script failed to fetch CC licenses that may be recorded as hyphenated variants like (CC-BY, CC-BY-NC, etc). I have implemented a compiled regex pattern that replaces the string matching for more robust license detection

I have also looked at some of the implementations in other PR to use the normalize_license_text() function for consistent license identification.

Please i'ld like to know what your thoughts are on these changes and work continuously on further improvements, Thanks.

It's probably a good idea to create a function in the shared library eventually. Please leave that to last, however.

Refactor arxiv_fetch.py to use requests library for HTTP requests, implementing retry logic for better error handling. Update license extraction logic and CSV headers to remove PLAN_INDEX.
@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot Hello, are there any more changes you'll like to be made to the ones already identified and done. Also i'ld like to ask, if there are no more changes to be made on this, what steps would you recommend as the next one for me in the project, with respect to integrating ArXiv as a valid source for the project?

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Please update sources.md for arXiv

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git Please update sources.md for arXiv

@TimidRobot, Thanks, Im happy we've gotten to this point. very excited to continue contributing to the project. I would like your permission to continue working on the processing and reporting scripts for the arXiv data source

@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hello @TimidRobot as per #179 (comment)
the implementation of the bucketing for the AUTHOR_COUNT was because you suggested that it was not so insightful to just merely dump the data, and so i thought a sort of grouping of this data would help us to generate meaningful insights. Raw author counts would create extremely sparse data with many single-occurrence values while Bucketing ("1", "2-3", "4-6", "7-10", "11+") on the other hand, creates meaningful statistical groups for analysis. It also facilitates visualization and reporting by creating compact, aggregated datasets that are easier to process and analyze hence supporting the project's goal of understanding "how knowledge and culture of the commons is distributed".

although these are my thoughts, I am still very happy to know what you think about this. To further share my opinion, these are some helpful links that guided my decision for bucketing the AUTHOR_COUNT data.

Scientometrics Journal Guidelines: (https://link.springer.com/journal/11192)
• Standard practice in bibliometric studies to group author counts into ranges for collaboration analysis which reduces noise and enables meaningful statistical comparisons across disciplines

FAIR Data Principles: (https://www.go-fair.org/fair-principles/)
• Bucketing supports Findability and Reusability by creating standardized categorical data
NIH Data Management Guidelines: (https://sharing.nih.gov/data-management-and-sharing-policy)

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Goziee-git please see two unresolved conversations:

  1. Add arXiv data fetching functionality #179 (comment)
  2. Add arXiv data fetching functionality #179 (comment)

@TimidRobot the conversations mentioned here are identical, please i'ld like to know what your preference is with the AUTHOR_COUNT, especially if you want it removed or modified in differnt way

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@TimidRobot

Copy link
Copy Markdown
Member

@Goziee-git Sorry, let me clarify. I think bucketing authors into AUTHOR_BUCKET is a good and helpful course of action. I wouldd also add https://en.wikipedia.org/wiki/Data_binning to the list of references.

I think which buckets are selected could use some refinement. My thinking is that there are three considerations:

  1. plotting
  2. precedence
  3. distribution

Plotting

For plotting, I think around five values display well.

Precedence

For precedence, we can look at the various citation styles:

Distribution

The script can be modified to provide information on all relevant author counts (only results above 1% are shown):

AuthorsCountPercent
313122.55%
211119.10%
49215.83%
18514.63%
5467.92%
6233.96%
7213.61%
8193.27%
9152.58%
10101.72%

Recommendation

I recommend the following buckets:

  • 1 author
  • 2 authors
  • 3 authors
  • 4 authors
  • 5+ authors

Comment threadscripts/1-fetch/arxiv_fetch.py Outdated
@Goziee-git

Goziee-git commented Oct 31, 2025

Copy link
Copy Markdown
ContributorAuthor

@TimidRobot as per ordering the constants, what would be your preference here. The current order of the constants is:

  1. FILE_ARXIV_COUNT (arxiv_1_count.csv)
  2. FILE_ARXIV_CATEGORY_REPORT (arxiv_2_count_by_category_report.csv)
  3. FILE_ARXIV_YEAR (arxiv_3_count_by_year.csv)
  4. FILE_ARXIV_AUTHOR_BUCKET (arxiv_4_count_by_author_bucket.csv)

This follows the logical/sequential order based on the numbered workflow (1 → 2 → 3 → 4).

in comparison to the gcs_fetch.py (current order):

  1. FILE1_COUNT (gcs_1_count.csv)
  2. FILE2_LANGUAGE (gcs_2_count_by_language.csv)
  3. FILE3_COUNTRY (gcs_3_count_by_country.csv)

Comparison:

  • Both scripts follow the same logical/sequential ordering pattern
  • Both order constants by their numbered workflow sequence (1 → 2 → 3 → 4)
  • Both use the numbering in the filename to determine order

Do you prefer some other order for this, so i can be guided as to the best course of action here

@Babi-B

Babi-B commented Oct 31, 2025

Copy link
Copy Markdown
Contributor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before
FILE_ARXIV_CATEGORY_REPORT

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

Hi @Goziee-git !

Order or sort alphabetically.

FILE_ARXIV_COUNT (arxiv_1_count.csv) shouldn't come before FILE_ARXIV_CATEGORY_REPORT

@Babi-B, Thanks for the clarity, i appreciate your reviews here.

@Goziee-git

Copy link
Copy Markdown
ContributorAuthor

@Babi-B, as per #185 looks like you've not made additions to sources.md. A moment ago i tried to read through and no links or references. I think you should update that. Or would you like me to go ahead with that 🚀🚀🚀🚀🚀

@TimidRobot

Copy link
Copy Markdown
Member

Both use the numbering in the filename to determine order

@Goziee-git The difference is that gcs_fetch.py uses a number in the constant so that they follow the flow when ordered:

FILE1_COUNT
^
FILE2_LANGUAGE
^

@TimidRobotTimidRobot left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great work, thank you!!

@TimidRobot
TimidRobot merged commit 0d44547 into creativecommons:mainNov 1, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Archived in project

Development

Successfully merging this pull request may close these issues.

Integrate arXiv as data source for academic commons quantification

4 participants

@Goziee-git@TimidRobot@Babi-B@cc-open-source-bot