Skip to content

Repository files navigation

Legal Feature Enhanced Semantic Matching Network for Similar Case Matching

Description

This repository is the source code of the paper Legal Feature Enhanced Semantic Matching Network for Similar Case Matching implemented via PyTorch.

Model Overview

ModelFig. 1 Overview of LFESM

Install and Run

Install

  • Python 3.6+

  • PyTorch 1.1.0+

  • Python requirements: run pip install -r requirements.txt.

  • Nvidia Apex (optional): Nvidia Apex enables mixed-precision training to accelerate the training procedure and decrease the memory usage. See official doc for installation. Specify fp16 = False in train.py to disable it.

  • Hardware: we recommend to use GPU to train LFESM. In our experiment, when we train the model with batch_size = 3 and fp16 = True on 2* GeForce RTX 2080, it takes 1~1.5 hour to finish one epoch.

  • Dataset: see Dataset.

  • BERT pretrained model: download the pretrained model here, and unzip the model into ./bert folder. See OpenCLaP for more details.

Train

python train.py

Our default parameters:

config= {
"max_length": 512,
"epochs": 6,
"batch_size": 3,
"learning_rate": 2e-5,
"fp16": True,
"fp16_opt_level": "O1",
"max_grad_norm": 1.0,
"warmup_steps": 0.1,
}

Predict

python predict.py

The output of the prediction is stored in ./data/test/output.txt.

Run ./scripts/judger.py to calculate the accuracy score.

Dataset

Download the dataset CAIL2019-SCM from here. Check CAIL2019 for more details about the dataset.

Table 1: The amount of data in CAIL2019-SCM

Datasetsim(a, b)>sim(a,c)​sim(a,b)<sim(a,c)​Total Amount
Train2,5962,5065,102
Valid8376631,500
Test8037331,536

Unzip the dataset and put the train, valid, and test set into raw, valid, and test folder of ./data folder.

Data Augmentation

To fulfill the distribution of dataset and enhance the performance of model training, we apply data augmentation to the dataset.

Let's denote the original triplet as (A, B, C). We add (A, C, B) into the dataset, which makes the amount multiply twice. We also tried other methods like (B, C, A) and (B, A, C), but they did not work.

Project Files

lfesm
├── bert # BERT pretrained model
├── config.py # Model config and hyper parameter
├── data # Store the dataset
├── data.py # Define the dataset
│ └── ...
├── model.py # Define the model trainer
├── models # Define the models
│ ├── baseline
│ ├── esim
│ ├── feature.py
│ ├── lfesm.py
├── predict.py # Predict
├── scripts # Utilities
│ └── ...
├── train.py # Train
└── util.py # Utility function

Results

Table 2: Experimental results of methods on CAIL2019-SCM

MethodValidTest
BaselineBERT61.9367.32
LSTM62.0068.00
CNN62.2769.53
Our BaselineBERT64.5365.59
LSTM64.3366.34
CNN64.7367.25
Best Score11.2yuan66.7372.07
backward67.7371.81
AlphaCourt70.0772.66
Our MethodLFESM70.0174.15

Reference

[1] coetaur0/ESIM

[2] padeoe/cail2019

[3] thunlp/OpenCLaP

[4] CAIL2019-SCM

[5] Taoooo9/Cail_Text_similarity_esimtribert

Acknowledgement

We sincerely appreciate Taoooo9‘s help.


Author: Zhilong Hong, Qifei Zhou, Rong Zhang, Weiping Li, and Tong Mo.

About

Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching".

Topics

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - Thesharing/LFESM: Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching". · GitHub
Skip to content

Repository files navigation

Legal Feature Enhanced Semantic Matching Network for Similar Case Matching

Description

This repository is the source code of the paper Legal Feature Enhanced Semantic Matching Network for Similar Case Matching implemented via PyTorch.

Model Overview

ModelFig. 1 Overview of LFESM

Install and Run

Install

  • Python 3.6+

  • PyTorch 1.1.0+

  • Python requirements: run pip install -r requirements.txt.

  • Nvidia Apex (optional): Nvidia Apex enables mixed-precision training to accelerate the training procedure and decrease the memory usage. See official doc for installation. Specify fp16 = False in train.py to disable it.

  • Hardware: we recommend to use GPU to train LFESM. In our experiment, when we train the model with batch_size = 3 and fp16 = True on 2* GeForce RTX 2080, it takes 1~1.5 hour to finish one epoch.

  • Dataset: see Dataset.

  • BERT pretrained model: download the pretrained model here, and unzip the model into ./bert folder. See OpenCLaP for more details.

Train

python train.py

Our default parameters:

config= {
"max_length": 512,
"epochs": 6,
"batch_size": 3,
"learning_rate": 2e-5,
"fp16": True,
"fp16_opt_level": "O1",
"max_grad_norm": 1.0,
"warmup_steps": 0.1,
}

Predict

python predict.py

The output of the prediction is stored in ./data/test/output.txt.

Run ./scripts/judger.py to calculate the accuracy score.

Dataset

Download the dataset CAIL2019-SCM from here. Check CAIL2019 for more details about the dataset.

Table 1: The amount of data in CAIL2019-SCM

Datasetsim(a, b)>sim(a,c)​sim(a,b)<sim(a,c)​Total Amount
Train2,5962,5065,102
Valid8376631,500
Test8037331,536

Unzip the dataset and put the train, valid, and test set into raw, valid, and test folder of ./data folder.

Data Augmentation

To fulfill the distribution of dataset and enhance the performance of model training, we apply data augmentation to the dataset.

Let's denote the original triplet as (A, B, C). We add (A, C, B) into the dataset, which makes the amount multiply twice. We also tried other methods like (B, C, A) and (B, A, C), but they did not work.

Project Files

lfesm
├── bert # BERT pretrained model
├── config.py # Model config and hyper parameter
├── data # Store the dataset
├── data.py # Define the dataset
│ └── ...
├── model.py # Define the model trainer
├── models # Define the models
│ ├── baseline
│ ├── esim
│ ├── feature.py
│ ├── lfesm.py
├── predict.py # Predict
├── scripts # Utilities
│ └── ...
├── train.py # Train
└── util.py # Utility function

Results

Table 2: Experimental results of methods on CAIL2019-SCM

MethodValidTest
BaselineBERT61.9367.32
LSTM62.0068.00
CNN62.2769.53
Our BaselineBERT64.5365.59
LSTM64.3366.34
CNN64.7367.25
Best Score11.2yuan66.7372.07
backward67.7371.81
AlphaCourt70.0772.66
Our MethodLFESM70.0174.15

Reference

[1] coetaur0/ESIM

[2] padeoe/cail2019

[3] thunlp/OpenCLaP

[4] CAIL2019-SCM

[5] Taoooo9/Cail_Text_similarity_esimtribert

Acknowledgement

We sincerely appreciate Taoooo9‘s help.


Author: Zhilong Hong, Qifei Zhou, Rong Zhang, Weiping Li, and Tong Mo.

About

Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching".

Topics

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - Thesharing/LFESM: Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching". · GitHub
Skip to content

Repository files navigation

Legal Feature Enhanced Semantic Matching Network for Similar Case Matching

Description

This repository is the source code of the paper Legal Feature Enhanced Semantic Matching Network for Similar Case Matching implemented via PyTorch.

Model Overview

ModelFig. 1 Overview of LFESM

Install and Run

Install

  • Python 3.6+

  • PyTorch 1.1.0+

  • Python requirements: run pip install -r requirements.txt.

  • Nvidia Apex (optional): Nvidia Apex enables mixed-precision training to accelerate the training procedure and decrease the memory usage. See official doc for installation. Specify fp16 = False in train.py to disable it.

  • Hardware: we recommend to use GPU to train LFESM. In our experiment, when we train the model with batch_size = 3 and fp16 = True on 2* GeForce RTX 2080, it takes 1~1.5 hour to finish one epoch.

  • Dataset: see Dataset.

  • BERT pretrained model: download the pretrained model here, and unzip the model into ./bert folder. See OpenCLaP for more details.

Train

python train.py

Our default parameters:

config= {
"max_length": 512,
"epochs": 6,
"batch_size": 3,
"learning_rate": 2e-5,
"fp16": True,
"fp16_opt_level": "O1",
"max_grad_norm": 1.0,
"warmup_steps": 0.1,
}

Predict

python predict.py

The output of the prediction is stored in ./data/test/output.txt.

Run ./scripts/judger.py to calculate the accuracy score.

Dataset

Download the dataset CAIL2019-SCM from here. Check CAIL2019 for more details about the dataset.

Table 1: The amount of data in CAIL2019-SCM

Datasetsim(a, b)>sim(a,c)​sim(a,b)<sim(a,c)​Total Amount
Train2,5962,5065,102
Valid8376631,500
Test8037331,536

Unzip the dataset and put the train, valid, and test set into raw, valid, and test folder of ./data folder.

Data Augmentation

To fulfill the distribution of dataset and enhance the performance of model training, we apply data augmentation to the dataset.

Let's denote the original triplet as (A, B, C). We add (A, C, B) into the dataset, which makes the amount multiply twice. We also tried other methods like (B, C, A) and (B, A, C), but they did not work.

Project Files

lfesm
├── bert # BERT pretrained model
├── config.py # Model config and hyper parameter
├── data # Store the dataset
├── data.py # Define the dataset
│ └── ...
├── model.py # Define the model trainer
├── models # Define the models
│ ├── baseline
│ ├── esim
│ ├── feature.py
│ ├── lfesm.py
├── predict.py # Predict
├── scripts # Utilities
│ └── ...
├── train.py # Train
└── util.py # Utility function

Results

Table 2: Experimental results of methods on CAIL2019-SCM

MethodValidTest
BaselineBERT61.9367.32
LSTM62.0068.00
CNN62.2769.53
Our BaselineBERT64.5365.59
LSTM64.3366.34
CNN64.7367.25
Best Score11.2yuan66.7372.07
backward67.7371.81
AlphaCourt70.0772.66
Our MethodLFESM70.0174.15

Reference

[1] coetaur0/ESIM

[2] padeoe/cail2019

[3] thunlp/OpenCLaP

[4] CAIL2019-SCM

[5] Taoooo9/Cail_Text_similarity_esimtribert

Acknowledgement

We sincerely appreciate Taoooo9‘s help.


Author: Zhilong Hong, Qifei Zhou, Rong Zhang, Weiping Li, and Tong Mo.

About

Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching".

Topics

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - Thesharing/LFESM: Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching". · GitHub
Skip to content

Repository files navigation

Legal Feature Enhanced Semantic Matching Network for Similar Case Matching

Description

This repository is the source code of the paper Legal Feature Enhanced Semantic Matching Network for Similar Case Matching implemented via PyTorch.

Model Overview

ModelFig. 1 Overview of LFESM

Install and Run

Install

  • Python 3.6+

  • PyTorch 1.1.0+

  • Python requirements: run pip install -r requirements.txt.

  • Nvidia Apex (optional): Nvidia Apex enables mixed-precision training to accelerate the training procedure and decrease the memory usage. See official doc for installation. Specify fp16 = False in train.py to disable it.

  • Hardware: we recommend to use GPU to train LFESM. In our experiment, when we train the model with batch_size = 3 and fp16 = True on 2* GeForce RTX 2080, it takes 1~1.5 hour to finish one epoch.

  • Dataset: see Dataset.

  • BERT pretrained model: download the pretrained model here, and unzip the model into ./bert folder. See OpenCLaP for more details.

Train

python train.py

Our default parameters:

config= {
"max_length": 512,
"epochs": 6,
"batch_size": 3,
"learning_rate": 2e-5,
"fp16": True,
"fp16_opt_level": "O1",
"max_grad_norm": 1.0,
"warmup_steps": 0.1,
}

Predict

python predict.py

The output of the prediction is stored in ./data/test/output.txt.

Run ./scripts/judger.py to calculate the accuracy score.

Dataset

Download the dataset CAIL2019-SCM from here. Check CAIL2019 for more details about the dataset.

Table 1: The amount of data in CAIL2019-SCM

Datasetsim(a, b)>sim(a,c)​sim(a,b)<sim(a,c)​Total Amount
Train2,5962,5065,102
Valid8376631,500
Test8037331,536

Unzip the dataset and put the train, valid, and test set into raw, valid, and test folder of ./data folder.

Data Augmentation

To fulfill the distribution of dataset and enhance the performance of model training, we apply data augmentation to the dataset.

Let's denote the original triplet as (A, B, C). We add (A, C, B) into the dataset, which makes the amount multiply twice. We also tried other methods like (B, C, A) and (B, A, C), but they did not work.

Project Files

lfesm
├── bert # BERT pretrained model
├── config.py # Model config and hyper parameter
├── data # Store the dataset
├── data.py # Define the dataset
│ └── ...
├── model.py # Define the model trainer
├── models # Define the models
│ ├── baseline
│ ├── esim
│ ├── feature.py
│ ├── lfesm.py
├── predict.py # Predict
├── scripts # Utilities
│ └── ...
├── train.py # Train
└── util.py # Utility function

Results

Table 2: Experimental results of methods on CAIL2019-SCM

MethodValidTest
BaselineBERT61.9367.32
LSTM62.0068.00
CNN62.2769.53
Our BaselineBERT64.5365.59
LSTM64.3366.34
CNN64.7367.25
Best Score11.2yuan66.7372.07
backward67.7371.81
AlphaCourt70.0772.66
Our MethodLFESM70.0174.15

Reference

[1] coetaur0/ESIM

[2] padeoe/cail2019

[3] thunlp/OpenCLaP

[4] CAIL2019-SCM

[5] Taoooo9/Cail_Text_similarity_esimtribert

Acknowledgement

We sincerely appreciate Taoooo9‘s help.


Author: Zhilong Hong, Qifei Zhou, Rong Zhang, Weiping Li, and Tong Mo.

About

Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching".

Topics

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - Thesharing/LFESM: Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching". · GitHub
Skip to content

Repository files navigation

Legal Feature Enhanced Semantic Matching Network for Similar Case Matching

Description

This repository is the source code of the paper Legal Feature Enhanced Semantic Matching Network for Similar Case Matching implemented via PyTorch.

Model Overview

ModelFig. 1 Overview of LFESM

Install and Run

Install

  • Python 3.6+

  • PyTorch 1.1.0+

  • Python requirements: run pip install -r requirements.txt.

  • Nvidia Apex (optional): Nvidia Apex enables mixed-precision training to accelerate the training procedure and decrease the memory usage. See official doc for installation. Specify fp16 = False in train.py to disable it.

  • Hardware: we recommend to use GPU to train LFESM. In our experiment, when we train the model with batch_size = 3 and fp16 = True on 2* GeForce RTX 2080, it takes 1~1.5 hour to finish one epoch.

  • Dataset: see Dataset.

  • BERT pretrained model: download the pretrained model here, and unzip the model into ./bert folder. See OpenCLaP for more details.

Train

python train.py

Our default parameters:

config= {
"max_length": 512,
"epochs": 6,
"batch_size": 3,
"learning_rate": 2e-5,
"fp16": True,
"fp16_opt_level": "O1",
"max_grad_norm": 1.0,
"warmup_steps": 0.1,
}

Predict

python predict.py

The output of the prediction is stored in ./data/test/output.txt.

Run ./scripts/judger.py to calculate the accuracy score.

Dataset

Download the dataset CAIL2019-SCM from here. Check CAIL2019 for more details about the dataset.

Table 1: The amount of data in CAIL2019-SCM

Datasetsim(a, b)>sim(a,c)​sim(a,b)<sim(a,c)​Total Amount
Train2,5962,5065,102
Valid8376631,500
Test8037331,536

Unzip the dataset and put the train, valid, and test set into raw, valid, and test folder of ./data folder.

Data Augmentation

To fulfill the distribution of dataset and enhance the performance of model training, we apply data augmentation to the dataset.

Let's denote the original triplet as (A, B, C). We add (A, C, B) into the dataset, which makes the amount multiply twice. We also tried other methods like (B, C, A) and (B, A, C), but they did not work.

Project Files

lfesm
├── bert # BERT pretrained model
├── config.py # Model config and hyper parameter
├── data # Store the dataset
├── data.py # Define the dataset
│ └── ...
├── model.py # Define the model trainer
├── models # Define the models
│ ├── baseline
│ ├── esim
│ ├── feature.py
│ ├── lfesm.py
├── predict.py # Predict
├── scripts # Utilities
│ └── ...
├── train.py # Train
└── util.py # Utility function

Results

Table 2: Experimental results of methods on CAIL2019-SCM

MethodValidTest
BaselineBERT61.9367.32
LSTM62.0068.00
CNN62.2769.53
Our BaselineBERT64.5365.59
LSTM64.3366.34
CNN64.7367.25
Best Score11.2yuan66.7372.07
backward67.7371.81
AlphaCourt70.0772.66
Our MethodLFESM70.0174.15

Reference

[1] coetaur0/ESIM

[2] padeoe/cail2019

[3] thunlp/OpenCLaP

[4] CAIL2019-SCM

[5] Taoooo9/Cail_Text_similarity_esimtribert

Acknowledgement

We sincerely appreciate Taoooo9‘s help.


Author: Zhilong Hong, Qifei Zhou, Rong Zhang, Weiping Li, and Tong Mo.

About

Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching".

Topics

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - Thesharing/LFESM: Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching". · GitHub
Skip to content

Repository files navigation

Legal Feature Enhanced Semantic Matching Network for Similar Case Matching

Description

This repository is the source code of the paper Legal Feature Enhanced Semantic Matching Network for Similar Case Matching implemented via PyTorch.

Model Overview

ModelFig. 1 Overview of LFESM

Install and Run

Install

  • Python 3.6+

  • PyTorch 1.1.0+

  • Python requirements: run pip install -r requirements.txt.

  • Nvidia Apex (optional): Nvidia Apex enables mixed-precision training to accelerate the training procedure and decrease the memory usage. See official doc for installation. Specify fp16 = False in train.py to disable it.

  • Hardware: we recommend to use GPU to train LFESM. In our experiment, when we train the model with batch_size = 3 and fp16 = True on 2* GeForce RTX 2080, it takes 1~1.5 hour to finish one epoch.

  • Dataset: see Dataset.

  • BERT pretrained model: download the pretrained model here, and unzip the model into ./bert folder. See OpenCLaP for more details.

Train

python train.py

Our default parameters:

config= {
"max_length": 512,
"epochs": 6,
"batch_size": 3,
"learning_rate": 2e-5,
"fp16": True,
"fp16_opt_level": "O1",
"max_grad_norm": 1.0,
"warmup_steps": 0.1,
}

Predict

python predict.py

The output of the prediction is stored in ./data/test/output.txt.

Run ./scripts/judger.py to calculate the accuracy score.

Dataset

Download the dataset CAIL2019-SCM from here. Check CAIL2019 for more details about the dataset.

Table 1: The amount of data in CAIL2019-SCM

Datasetsim(a, b)>sim(a,c)​sim(a,b)<sim(a,c)​Total Amount
Train2,5962,5065,102
Valid8376631,500
Test8037331,536

Unzip the dataset and put the train, valid, and test set into raw, valid, and test folder of ./data folder.

Data Augmentation

To fulfill the distribution of dataset and enhance the performance of model training, we apply data augmentation to the dataset.

Let's denote the original triplet as (A, B, C). We add (A, C, B) into the dataset, which makes the amount multiply twice. We also tried other methods like (B, C, A) and (B, A, C), but they did not work.

Project Files

lfesm
├── bert # BERT pretrained model
├── config.py # Model config and hyper parameter
├── data # Store the dataset
├── data.py # Define the dataset
│ └── ...
├── model.py # Define the model trainer
├── models # Define the models
│ ├── baseline
│ ├── esim
│ ├── feature.py
│ ├── lfesm.py
├── predict.py # Predict
├── scripts # Utilities
│ └── ...
├── train.py # Train
└── util.py # Utility function

Results

Table 2: Experimental results of methods on CAIL2019-SCM

MethodValidTest
BaselineBERT61.9367.32
LSTM62.0068.00
CNN62.2769.53
Our BaselineBERT64.5365.59
LSTM64.3366.34
CNN64.7367.25
Best Score11.2yuan66.7372.07
backward67.7371.81
AlphaCourt70.0772.66
Our MethodLFESM70.0174.15

Reference

[1] coetaur0/ESIM

[2] padeoe/cail2019

[3] thunlp/OpenCLaP

[4] CAIL2019-SCM

[5] Taoooo9/Cail_Text_similarity_esimtribert

Acknowledgement

We sincerely appreciate Taoooo9‘s help.


Author: Zhilong Hong, Qifei Zhou, Rong Zhang, Weiping Li, and Tong Mo.

About

Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching".

Topics

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - Thesharing/LFESM: Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching". · GitHub
Skip to content

Repository files navigation

Legal Feature Enhanced Semantic Matching Network for Similar Case Matching

Description

This repository is the source code of the paper Legal Feature Enhanced Semantic Matching Network for Similar Case Matching implemented via PyTorch.

Model Overview

ModelFig. 1 Overview of LFESM

Install and Run

Install

  • Python 3.6+

  • PyTorch 1.1.0+

  • Python requirements: run pip install -r requirements.txt.

  • Nvidia Apex (optional): Nvidia Apex enables mixed-precision training to accelerate the training procedure and decrease the memory usage. See official doc for installation. Specify fp16 = False in train.py to disable it.

  • Hardware: we recommend to use GPU to train LFESM. In our experiment, when we train the model with batch_size = 3 and fp16 = True on 2* GeForce RTX 2080, it takes 1~1.5 hour to finish one epoch.

  • Dataset: see Dataset.

  • BERT pretrained model: download the pretrained model here, and unzip the model into ./bert folder. See OpenCLaP for more details.

Train

python train.py

Our default parameters:

config= {
"max_length": 512,
"epochs": 6,
"batch_size": 3,
"learning_rate": 2e-5,
"fp16": True,
"fp16_opt_level": "O1",
"max_grad_norm": 1.0,
"warmup_steps": 0.1,
}

Predict

python predict.py

The output of the prediction is stored in ./data/test/output.txt.

Run ./scripts/judger.py to calculate the accuracy score.

Dataset

Download the dataset CAIL2019-SCM from here. Check CAIL2019 for more details about the dataset.

Table 1: The amount of data in CAIL2019-SCM

Datasetsim(a, b)>sim(a,c)​sim(a,b)<sim(a,c)​Total Amount
Train2,5962,5065,102
Valid8376631,500
Test8037331,536

Unzip the dataset and put the train, valid, and test set into raw, valid, and test folder of ./data folder.

Data Augmentation

To fulfill the distribution of dataset and enhance the performance of model training, we apply data augmentation to the dataset.

Let's denote the original triplet as (A, B, C). We add (A, C, B) into the dataset, which makes the amount multiply twice. We also tried other methods like (B, C, A) and (B, A, C), but they did not work.

Project Files

lfesm
├── bert # BERT pretrained model
├── config.py # Model config and hyper parameter
├── data # Store the dataset
├── data.py # Define the dataset
│ └── ...
├── model.py # Define the model trainer
├── models # Define the models
│ ├── baseline
│ ├── esim
│ ├── feature.py
│ ├── lfesm.py
├── predict.py # Predict
├── scripts # Utilities
│ └── ...
├── train.py # Train
└── util.py # Utility function

Results

Table 2: Experimental results of methods on CAIL2019-SCM

MethodValidTest
BaselineBERT61.9367.32
LSTM62.0068.00
CNN62.2769.53
Our BaselineBERT64.5365.59
LSTM64.3366.34
CNN64.7367.25
Best Score11.2yuan66.7372.07
backward67.7371.81
AlphaCourt70.0772.66
Our MethodLFESM70.0174.15

Reference

[1] coetaur0/ESIM

[2] padeoe/cail2019

[3] thunlp/OpenCLaP

[4] CAIL2019-SCM

[5] Taoooo9/Cail_Text_similarity_esimtribert

Acknowledgement

We sincerely appreciate Taoooo9‘s help.


Author: Zhilong Hong, Qifei Zhou, Rong Zhang, Weiping Li, and Tong Mo.

About

Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching".

Topics

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - Thesharing/LFESM: Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching". · GitHub
Skip to content

Repository files navigation

Legal Feature Enhanced Semantic Matching Network for Similar Case Matching

Description

This repository is the source code of the paper Legal Feature Enhanced Semantic Matching Network for Similar Case Matching implemented via PyTorch.

Model Overview

ModelFig. 1 Overview of LFESM

Install and Run

Install

  • Python 3.6+

  • PyTorch 1.1.0+

  • Python requirements: run pip install -r requirements.txt.

  • Nvidia Apex (optional): Nvidia Apex enables mixed-precision training to accelerate the training procedure and decrease the memory usage. See official doc for installation. Specify fp16 = False in train.py to disable it.

  • Hardware: we recommend to use GPU to train LFESM. In our experiment, when we train the model with batch_size = 3 and fp16 = True on 2* GeForce RTX 2080, it takes 1~1.5 hour to finish one epoch.

  • Dataset: see Dataset.

  • BERT pretrained model: download the pretrained model here, and unzip the model into ./bert folder. See OpenCLaP for more details.

Train

python train.py

Our default parameters:

config= {
"max_length": 512,
"epochs": 6,
"batch_size": 3,
"learning_rate": 2e-5,
"fp16": True,
"fp16_opt_level": "O1",
"max_grad_norm": 1.0,
"warmup_steps": 0.1,
}

Predict

python predict.py

The output of the prediction is stored in ./data/test/output.txt.

Run ./scripts/judger.py to calculate the accuracy score.

Dataset

Download the dataset CAIL2019-SCM from here. Check CAIL2019 for more details about the dataset.

Table 1: The amount of data in CAIL2019-SCM

Datasetsim(a, b)>sim(a,c)​sim(a,b)<sim(a,c)​Total Amount
Train2,5962,5065,102
Valid8376631,500
Test8037331,536

Unzip the dataset and put the train, valid, and test set into raw, valid, and test folder of ./data folder.

Data Augmentation

To fulfill the distribution of dataset and enhance the performance of model training, we apply data augmentation to the dataset.

Let's denote the original triplet as (A, B, C). We add (A, C, B) into the dataset, which makes the amount multiply twice. We also tried other methods like (B, C, A) and (B, A, C), but they did not work.

Project Files

lfesm
├── bert # BERT pretrained model
├── config.py # Model config and hyper parameter
├── data # Store the dataset
├── data.py # Define the dataset
│ └── ...
├── model.py # Define the model trainer
├── models # Define the models
│ ├── baseline
│ ├── esim
│ ├── feature.py
│ ├── lfesm.py
├── predict.py # Predict
├── scripts # Utilities
│ └── ...
├── train.py # Train
└── util.py # Utility function

Results

Table 2: Experimental results of methods on CAIL2019-SCM

MethodValidTest
BaselineBERT61.9367.32
LSTM62.0068.00
CNN62.2769.53
Our BaselineBERT64.5365.59
LSTM64.3366.34
CNN64.7367.25
Best Score11.2yuan66.7372.07
backward67.7371.81
AlphaCourt70.0772.66
Our MethodLFESM70.0174.15

Reference

[1] coetaur0/ESIM

[2] padeoe/cail2019

[3] thunlp/OpenCLaP

[4] CAIL2019-SCM

[5] Taoooo9/Cail_Text_similarity_esimtribert

Acknowledgement

We sincerely appreciate Taoooo9‘s help.


Author: Zhilong Hong, Qifei Zhou, Rong Zhang, Weiping Li, and Tong Mo.

About

Source code for Paper "Legal Feature Enhanced Semantic Matching Network for Similar Case Matching".

Topics

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages