Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Reference Language Corpora (RLC)

This repository contains programming language grammars and example programs in those languages.

This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Subdirectories

languages contains grammars and example inputs for well some well-known general purpose languages. The rhul subfolders contain examples that we have used in many of our research papers. For some languages we have also included large sets of examples drawn from publicly available repositories: we thank the owners for making the material available. See the README.md file in individual directories for details on how these corpora were constructed.

experiments contains scripts for running experiments using our own parser implementations, and data collected from those experiments.

'org', 'cws' and 'tok' variants

The example inputs come in three flavours: the original strings as produced by the programmer (org), compressed whitespace versions (cws) and tokenised versions(tok).

org inputs are just normal example programs.

cws inputs contain the same strings as the corresponding org input except that every run of whitespace is replaced by a single space unless the run contains one or more newline characters, in which case the run is replaced by a single newline. The purpose of the cws variants is to normalise lexer performance so that long runs of whitespace do not introduce lexicalisation overheads for experiments that focus on parser performance.

tok inputs have keyword-only lexical structure in which lexical elements whose pattern contains more than one lexeme are replaced by a keyword. This means, for instance, that all identifiers in a Java program are replaced by the keyword ID, all integer constants by the keyword INTEGER, all string constants by the keyword STRING and so on.

Running experiments

experiments/RunExp.java scans the entire languages directory structure, running experiments specified in experiments/try and producing log.csv, timeSummary.csv and spaceSummary.csv

Directory structure

0 RLC root


1 .. experiments

2 .... scripts (a collection of ART and GTB stub scripts that are concatenated with grammars)

2 .... results (results from some of our experimental runs)

2 .... RunExp.java(Scan directory stucture and run experiments)

2 .... try(the subset of scripts to be used in the next run)


1 .. languages

2 .... (more languages...)


2 .... java

3 ...... build ( language standards and programs to construct the corpora )

3 ...... grammar

4 ........ (more grammars for java...)

4 ........ jls18 (grammars from the Java Language Specification version 18)

5 .......... org(grammar group for jls18 that include full lexical rules)

5 .......... tok(grammar group for jls18 that expect inputs to be 'tokenised')

3 ...... corpus

4 ........ (more corpora for Java...)

4 ........ jfx (a specific corpus for Java)

5 .......... org(the original version of each input)

5 .......... cws(compressed white space version of each input)

5 .......... tok('tokenised' input group)

Other resources

You may also find the following pages useful.

You will find source code and pre-built executables for GTB and the tokenisers that produce our tok variants in the old directory of the ART repository.

Adrian Johnstone, August 2024

About

This repository contains programming language grammars and example programs in those languages. This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Reference Language Corpora (RLC)

This repository contains programming language grammars and example programs in those languages.

This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Subdirectories

languages contains grammars and example inputs for well some well-known general purpose languages. The rhul subfolders contain examples that we have used in many of our research papers. For some languages we have also included large sets of examples drawn from publicly available repositories: we thank the owners for making the material available. See the README.md file in individual directories for details on how these corpora were constructed.

experiments contains scripts for running experiments using our own parser implementations, and data collected from those experiments.

'org', 'cws' and 'tok' variants

The example inputs come in three flavours: the original strings as produced by the programmer (org), compressed whitespace versions (cws) and tokenised versions(tok).

org inputs are just normal example programs.

cws inputs contain the same strings as the corresponding org input except that every run of whitespace is replaced by a single space unless the run contains one or more newline characters, in which case the run is replaced by a single newline. The purpose of the cws variants is to normalise lexer performance so that long runs of whitespace do not introduce lexicalisation overheads for experiments that focus on parser performance.

tok inputs have keyword-only lexical structure in which lexical elements whose pattern contains more than one lexeme are replaced by a keyword. This means, for instance, that all identifiers in a Java program are replaced by the keyword ID, all integer constants by the keyword INTEGER, all string constants by the keyword STRING and so on.

Running experiments

experiments/RunExp.java scans the entire languages directory structure, running experiments specified in experiments/try and producing log.csv, timeSummary.csv and spaceSummary.csv

Directory structure

0 RLC root


1 .. experiments

2 .... scripts (a collection of ART and GTB stub scripts that are concatenated with grammars)

2 .... results (results from some of our experimental runs)

2 .... RunExp.java(Scan directory stucture and run experiments)

2 .... try(the subset of scripts to be used in the next run)


1 .. languages

2 .... (more languages...)


2 .... java

3 ...... build ( language standards and programs to construct the corpora )

3 ...... grammar

4 ........ (more grammars for java...)

4 ........ jls18 (grammars from the Java Language Specification version 18)

5 .......... org(grammar group for jls18 that include full lexical rules)

5 .......... tok(grammar group for jls18 that expect inputs to be 'tokenised')

3 ...... corpus

4 ........ (more corpora for Java...)

4 ........ jfx (a specific corpus for Java)

5 .......... org(the original version of each input)

5 .......... cws(compressed white space version of each input)

5 .......... tok('tokenised' input group)

Other resources

You may also find the following pages useful.

You will find source code and pre-built executables for GTB and the tokenisers that produce our tok variants in the old directory of the ART repository.

Adrian Johnstone, August 2024

About

This repository contains programming language grammars and example programs in those languages. This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Reference Language Corpora (RLC)

This repository contains programming language grammars and example programs in those languages.

This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Subdirectories

languages contains grammars and example inputs for well some well-known general purpose languages. The rhul subfolders contain examples that we have used in many of our research papers. For some languages we have also included large sets of examples drawn from publicly available repositories: we thank the owners for making the material available. See the README.md file in individual directories for details on how these corpora were constructed.

experiments contains scripts for running experiments using our own parser implementations, and data collected from those experiments.

'org', 'cws' and 'tok' variants

The example inputs come in three flavours: the original strings as produced by the programmer (org), compressed whitespace versions (cws) and tokenised versions(tok).

org inputs are just normal example programs.

cws inputs contain the same strings as the corresponding org input except that every run of whitespace is replaced by a single space unless the run contains one or more newline characters, in which case the run is replaced by a single newline. The purpose of the cws variants is to normalise lexer performance so that long runs of whitespace do not introduce lexicalisation overheads for experiments that focus on parser performance.

tok inputs have keyword-only lexical structure in which lexical elements whose pattern contains more than one lexeme are replaced by a keyword. This means, for instance, that all identifiers in a Java program are replaced by the keyword ID, all integer constants by the keyword INTEGER, all string constants by the keyword STRING and so on.

Running experiments

experiments/RunExp.java scans the entire languages directory structure, running experiments specified in experiments/try and producing log.csv, timeSummary.csv and spaceSummary.csv

Directory structure

0 RLC root


1 .. experiments

2 .... scripts (a collection of ART and GTB stub scripts that are concatenated with grammars)

2 .... results (results from some of our experimental runs)

2 .... RunExp.java(Scan directory stucture and run experiments)

2 .... try(the subset of scripts to be used in the next run)


1 .. languages

2 .... (more languages...)


2 .... java

3 ...... build ( language standards and programs to construct the corpora )

3 ...... grammar

4 ........ (more grammars for java...)

4 ........ jls18 (grammars from the Java Language Specification version 18)

5 .......... org(grammar group for jls18 that include full lexical rules)

5 .......... tok(grammar group for jls18 that expect inputs to be 'tokenised')

3 ...... corpus

4 ........ (more corpora for Java...)

4 ........ jfx (a specific corpus for Java)

5 .......... org(the original version of each input)

5 .......... cws(compressed white space version of each input)

5 .......... tok('tokenised' input group)

Other resources

You may also find the following pages useful.

You will find source code and pre-built executables for GTB and the tokenisers that produce our tok variants in the old directory of the ART repository.

Adrian Johnstone, August 2024

About

This repository contains programming language grammars and example programs in those languages. This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Reference Language Corpora (RLC)

This repository contains programming language grammars and example programs in those languages.

This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Subdirectories

languages contains grammars and example inputs for well some well-known general purpose languages. The rhul subfolders contain examples that we have used in many of our research papers. For some languages we have also included large sets of examples drawn from publicly available repositories: we thank the owners for making the material available. See the README.md file in individual directories for details on how these corpora were constructed.

experiments contains scripts for running experiments using our own parser implementations, and data collected from those experiments.

'org', 'cws' and 'tok' variants

The example inputs come in three flavours: the original strings as produced by the programmer (org), compressed whitespace versions (cws) and tokenised versions(tok).

org inputs are just normal example programs.

cws inputs contain the same strings as the corresponding org input except that every run of whitespace is replaced by a single space unless the run contains one or more newline characters, in which case the run is replaced by a single newline. The purpose of the cws variants is to normalise lexer performance so that long runs of whitespace do not introduce lexicalisation overheads for experiments that focus on parser performance.

tok inputs have keyword-only lexical structure in which lexical elements whose pattern contains more than one lexeme are replaced by a keyword. This means, for instance, that all identifiers in a Java program are replaced by the keyword ID, all integer constants by the keyword INTEGER, all string constants by the keyword STRING and so on.

Running experiments

experiments/RunExp.java scans the entire languages directory structure, running experiments specified in experiments/try and producing log.csv, timeSummary.csv and spaceSummary.csv

Directory structure

0 RLC root


1 .. experiments

2 .... scripts (a collection of ART and GTB stub scripts that are concatenated with grammars)

2 .... results (results from some of our experimental runs)

2 .... RunExp.java(Scan directory stucture and run experiments)

2 .... try(the subset of scripts to be used in the next run)


1 .. languages

2 .... (more languages...)


2 .... java

3 ...... build ( language standards and programs to construct the corpora )

3 ...... grammar

4 ........ (more grammars for java...)

4 ........ jls18 (grammars from the Java Language Specification version 18)

5 .......... org(grammar group for jls18 that include full lexical rules)

5 .......... tok(grammar group for jls18 that expect inputs to be 'tokenised')

3 ...... corpus

4 ........ (more corpora for Java...)

4 ........ jfx (a specific corpus for Java)

5 .......... org(the original version of each input)

5 .......... cws(compressed white space version of each input)

5 .......... tok('tokenised' input group)

Other resources

You may also find the following pages useful.

You will find source code and pre-built executables for GTB and the tokenisers that produce our tok variants in the old directory of the ART repository.

Adrian Johnstone, August 2024

About

This repository contains programming language grammars and example programs in those languages. This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Reference Language Corpora (RLC)

This repository contains programming language grammars and example programs in those languages.

This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Subdirectories

languages contains grammars and example inputs for well some well-known general purpose languages. The rhul subfolders contain examples that we have used in many of our research papers. For some languages we have also included large sets of examples drawn from publicly available repositories: we thank the owners for making the material available. See the README.md file in individual directories for details on how these corpora were constructed.

experiments contains scripts for running experiments using our own parser implementations, and data collected from those experiments.

'org', 'cws' and 'tok' variants

The example inputs come in three flavours: the original strings as produced by the programmer (org), compressed whitespace versions (cws) and tokenised versions(tok).

org inputs are just normal example programs.

cws inputs contain the same strings as the corresponding org input except that every run of whitespace is replaced by a single space unless the run contains one or more newline characters, in which case the run is replaced by a single newline. The purpose of the cws variants is to normalise lexer performance so that long runs of whitespace do not introduce lexicalisation overheads for experiments that focus on parser performance.

tok inputs have keyword-only lexical structure in which lexical elements whose pattern contains more than one lexeme are replaced by a keyword. This means, for instance, that all identifiers in a Java program are replaced by the keyword ID, all integer constants by the keyword INTEGER, all string constants by the keyword STRING and so on.

Running experiments

experiments/RunExp.java scans the entire languages directory structure, running experiments specified in experiments/try and producing log.csv, timeSummary.csv and spaceSummary.csv

Directory structure

0 RLC root


1 .. experiments

2 .... scripts (a collection of ART and GTB stub scripts that are concatenated with grammars)

2 .... results (results from some of our experimental runs)

2 .... RunExp.java(Scan directory stucture and run experiments)

2 .... try(the subset of scripts to be used in the next run)


1 .. languages

2 .... (more languages...)


2 .... java

3 ...... build ( language standards and programs to construct the corpora )

3 ...... grammar

4 ........ (more grammars for java...)

4 ........ jls18 (grammars from the Java Language Specification version 18)

5 .......... org(grammar group for jls18 that include full lexical rules)

5 .......... tok(grammar group for jls18 that expect inputs to be 'tokenised')

3 ...... corpus

4 ........ (more corpora for Java...)

4 ........ jfx (a specific corpus for Java)

5 .......... org(the original version of each input)

5 .......... cws(compressed white space version of each input)

5 .......... tok('tokenised' input group)

Other resources

You may also find the following pages useful.

You will find source code and pre-built executables for GTB and the tokenisers that produce our tok variants in the old directory of the ART repository.

Adrian Johnstone, August 2024

About

This repository contains programming language grammars and example programs in those languages. This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Reference Language Corpora (RLC)

This repository contains programming language grammars and example programs in those languages.

This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Subdirectories

languages contains grammars and example inputs for well some well-known general purpose languages. The rhul subfolders contain examples that we have used in many of our research papers. For some languages we have also included large sets of examples drawn from publicly available repositories: we thank the owners for making the material available. See the README.md file in individual directories for details on how these corpora were constructed.

experiments contains scripts for running experiments using our own parser implementations, and data collected from those experiments.

'org', 'cws' and 'tok' variants

The example inputs come in three flavours: the original strings as produced by the programmer (org), compressed whitespace versions (cws) and tokenised versions(tok).

org inputs are just normal example programs.

cws inputs contain the same strings as the corresponding org input except that every run of whitespace is replaced by a single space unless the run contains one or more newline characters, in which case the run is replaced by a single newline. The purpose of the cws variants is to normalise lexer performance so that long runs of whitespace do not introduce lexicalisation overheads for experiments that focus on parser performance.

tok inputs have keyword-only lexical structure in which lexical elements whose pattern contains more than one lexeme are replaced by a keyword. This means, for instance, that all identifiers in a Java program are replaced by the keyword ID, all integer constants by the keyword INTEGER, all string constants by the keyword STRING and so on.

Running experiments

experiments/RunExp.java scans the entire languages directory structure, running experiments specified in experiments/try and producing log.csv, timeSummary.csv and spaceSummary.csv

Directory structure

0 RLC root


1 .. experiments

2 .... scripts (a collection of ART and GTB stub scripts that are concatenated with grammars)

2 .... results (results from some of our experimental runs)

2 .... RunExp.java(Scan directory stucture and run experiments)

2 .... try(the subset of scripts to be used in the next run)


1 .. languages

2 .... (more languages...)


2 .... java

3 ...... build ( language standards and programs to construct the corpora )

3 ...... grammar

4 ........ (more grammars for java...)

4 ........ jls18 (grammars from the Java Language Specification version 18)

5 .......... org(grammar group for jls18 that include full lexical rules)

5 .......... tok(grammar group for jls18 that expect inputs to be 'tokenised')

3 ...... corpus

4 ........ (more corpora for Java...)

4 ........ jfx (a specific corpus for Java)

5 .......... org(the original version of each input)

5 .......... cws(compressed white space version of each input)

5 .......... tok('tokenised' input group)

Other resources

You may also find the following pages useful.

You will find source code and pre-built executables for GTB and the tokenisers that produce our tok variants in the old directory of the ART repository.

Adrian Johnstone, August 2024

About

This repository contains programming language grammars and example programs in those languages. This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Reference Language Corpora (RLC)

This repository contains programming language grammars and example programs in those languages.

This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Subdirectories

languages contains grammars and example inputs for well some well-known general purpose languages. The rhul subfolders contain examples that we have used in many of our research papers. For some languages we have also included large sets of examples drawn from publicly available repositories: we thank the owners for making the material available. See the README.md file in individual directories for details on how these corpora were constructed.

experiments contains scripts for running experiments using our own parser implementations, and data collected from those experiments.

'org', 'cws' and 'tok' variants

The example inputs come in three flavours: the original strings as produced by the programmer (org), compressed whitespace versions (cws) and tokenised versions(tok).

org inputs are just normal example programs.

cws inputs contain the same strings as the corresponding org input except that every run of whitespace is replaced by a single space unless the run contains one or more newline characters, in which case the run is replaced by a single newline. The purpose of the cws variants is to normalise lexer performance so that long runs of whitespace do not introduce lexicalisation overheads for experiments that focus on parser performance.

tok inputs have keyword-only lexical structure in which lexical elements whose pattern contains more than one lexeme are replaced by a keyword. This means, for instance, that all identifiers in a Java program are replaced by the keyword ID, all integer constants by the keyword INTEGER, all string constants by the keyword STRING and so on.

Running experiments

experiments/RunExp.java scans the entire languages directory structure, running experiments specified in experiments/try and producing log.csv, timeSummary.csv and spaceSummary.csv

Directory structure

0 RLC root


1 .. experiments

2 .... scripts (a collection of ART and GTB stub scripts that are concatenated with grammars)

2 .... results (results from some of our experimental runs)

2 .... RunExp.java(Scan directory stucture and run experiments)

2 .... try(the subset of scripts to be used in the next run)


1 .. languages

2 .... (more languages...)


2 .... java

3 ...... build ( language standards and programs to construct the corpora )

3 ...... grammar

4 ........ (more grammars for java...)

4 ........ jls18 (grammars from the Java Language Specification version 18)

5 .......... org(grammar group for jls18 that include full lexical rules)

5 .......... tok(grammar group for jls18 that expect inputs to be 'tokenised')

3 ...... corpus

4 ........ (more corpora for Java...)

4 ........ jfx (a specific corpus for Java)

5 .......... org(the original version of each input)

5 .......... cws(compressed white space version of each input)

5 .......... tok('tokenised' input group)

Other resources

You may also find the following pages useful.

You will find source code and pre-built executables for GTB and the tokenisers that produce our tok variants in the old directory of the ART repository.

Adrian Johnstone, August 2024

About

This repository contains programming language grammars and example programs in those languages. This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

88 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Reference Language Corpora (RLC)

This repository contains programming language grammars and example programs in those languages.

This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Subdirectories

languages contains grammars and example inputs for well some well-known general purpose languages. The rhul subfolders contain examples that we have used in many of our research papers. For some languages we have also included large sets of examples drawn from publicly available repositories: we thank the owners for making the material available. See the README.md file in individual directories for details on how these corpora were constructed.

experiments contains scripts for running experiments using our own parser implementations, and data collected from those experiments.

'org', 'cws' and 'tok' variants

The example inputs come in three flavours: the original strings as produced by the programmer (org), compressed whitespace versions (cws) and tokenised versions(tok).

org inputs are just normal example programs.

cws inputs contain the same strings as the corresponding org input except that every run of whitespace is replaced by a single space unless the run contains one or more newline characters, in which case the run is replaced by a single newline. The purpose of the cws variants is to normalise lexer performance so that long runs of whitespace do not introduce lexicalisation overheads for experiments that focus on parser performance.

tok inputs have keyword-only lexical structure in which lexical elements whose pattern contains more than one lexeme are replaced by a keyword. This means, for instance, that all identifiers in a Java program are replaced by the keyword ID, all integer constants by the keyword INTEGER, all string constants by the keyword STRING and so on.

Running experiments

experiments/RunExp.java scans the entire languages directory structure, running experiments specified in experiments/try and producing log.csv, timeSummary.csv and spaceSummary.csv

Directory structure

0 RLC root


1 .. experiments

2 .... scripts (a collection of ART and GTB stub scripts that are concatenated with grammars)

2 .... results (results from some of our experimental runs)

2 .... RunExp.java(Scan directory stucture and run experiments)

2 .... try(the subset of scripts to be used in the next run)


1 .. languages

2 .... (more languages...)


2 .... java

3 ...... build ( language standards and programs to construct the corpora )

3 ...... grammar

4 ........ (more grammars for java...)

4 ........ jls18 (grammars from the Java Language Specification version 18)

5 .......... org(grammar group for jls18 that include full lexical rules)

5 .......... tok(grammar group for jls18 that expect inputs to be 'tokenised')

3 ...... corpus

4 ........ (more corpora for Java...)

4 ........ jfx (a specific corpus for Java)

5 .......... org(the original version of each input)

5 .......... cws(compressed white space version of each input)

5 .......... tok('tokenised' input group)

Other resources

You may also find the following pages useful.

You will find source code and pre-built executables for GTB and the tokenisers that produce our tok variants in the old directory of the ART repository.

Adrian Johnstone, August 2024

About

This repository contains programming language grammars and example programs in those languages. This collection is curated by Adrian Johnstone and Elizabeth Scott of the Centre for Software Language Engineering at Royal Holloway, University of London.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages