Latest commit

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Currency-Forecasting

Cumulative balance charts When forecasting in finance, derivatives of the price itself are usually used. Sometimes this has some effect, the improvements are barely noticeable, if not accidental. Famous expression: 'garbage in the input garbage out. You can try to look for signs not in the price itself. Central banks, various bureaus produce numerous macroeconomic indicators. I used them as additional features for forecasting currencies.

For each currency area, more than 700 types of data were obtained (for each currency a couple of about 1500). I tried to reduce the number of features using feature importance based on a random forest, the importance of permutations, etc. This did not give the desired effect. Probably because there is little information in the data. Then I started to add one feature to the main dataset and see if there are any improvements in the balance curve. So Thus, the number of signs was reduced tenfold. This job took a lot time. Then i began to cycle through these selected features adding three to the main dataset or applied a genetic algorithm to find the best combination. A variant with a genetic algorithm found the best combinations with a large number of features and over time it turned out that it was retraining (overfitting). In the above chart, the result obtained for EURUSD since 1991 (only sales).Sell and hold in orange. On the left, the model trained without adding macroeconomic indicators ('Depozit_no'), on the right with adding ('Depozit_n'). The model was trained in January 2020 and the settings have not changed since then. Classifier was used for classification: GradientBoostingClassifier. Signs news 1, 2, 3 do not plan to disclose.

To reproduce the code, the following libraries are needed: numpy, pandas, scipy, sklearn, matplotlib, seaborn, talib. You also need an include file: qnt.

By code:

1.The dataset.csv file is read, in which EURUSD daily quotes and three selected macroeconomic indicators (these are their momentums shifted by the required amount and split-normalized). Example: if the data is dated October 1st and becomes available on November 15th, then it is shifted by about 45 days.

2.The price series is converted into a series of technical analysis indicators. This data+ macroeconomic indicators create a training dataset.

3.Based on the classifier labels obtained, balance curves are built in dataframes df_n, df_no.

4.Various indicators, a graph of balances and distributions of transactions are displayed.

this is Depozit_no --- Pearson corr train 0.992 Pearson corr test 0.918
Sharp_Ratio_train 3.45 Sharp_Ratio_test 2.54 amount of deals 960
Train precision recall f1-score support
0 0.61 0.25 0.35 2807
1 0.54 0.84 0.66 2895
accuracy 0.55 5702
macro avg 0.57 0.55 0.51 5702
weighted avg 0.57 0.55 0.51 5702
Test precision recall f1-score support
0 0.56 0.17 0.27 1355
1 0.52 0.87 0.65 1383
accuracy 0.52 2738
macro avg 0.54 0.52 0.46 2738
weighted avg 0.54 0.52 0.46 2738
this is Depozit_n --- Pearson corr train 0.994 Pearson corr test 0.981
Sharp_Ratio_train 6.15 Sharp_Ratio_test 3.27 amount of deals 1027
Train precision recall f1-score support
0 0.73 0.30 0.43 2807
1 0.57 0.89 0.70 2895
accuracy 0.60 5702
macro avg 0.65 0.60 0.56 5702
weighted avg 0.65 0.60 0.56 5702
Test precision recall f1-score support
0 0.52 0.44 0.48 1355
1 0.53 0.61 0.56 1383
accuracy 0.52 2738
macro avg 0.52 0.52 0.52 2738
weighted avg 0.52 0.52 0.52 2738

distribution of deals

import scipy.stats
arr_no_ = arr_no[il[0]:]
arr_n_ = arr_n[il[1]:]
def ep(N, n):
p = n/N
sigma = math.sqrt(p * (1 - p)/N)
return p, sigma
def a_b_statistic(N_A, n_a, N_B, n_b):
p_A, sigma_A = ep(N_A, n_a)
p_B, sigma_B = ep(N_B, n_b)
return (p_B - p_A)/math.sqrt(sigma_A ** 2 + sigma_B ** 2)
z_csore = a_b_statistic(len(arr_n_), (arr_n_ > 0).sum(), len(arr_no_), (arr_no_ > 0).sum())
p_value = scipy.stats.norm.sf(abs(z_csore))
print('z_csore : ' + str(z_csore), 'p value : ' + str(p_value))
lno = round(arr_no_[arr_no_ < 0].sum(), 3)#amount of loss
pno = round(arr_no_[arr_no_ > 0].sum(), 3)#amount of profits
ano = round(abs(pno/lno), 3)#the ratio of the amount of profits to the amount of losses
ln = round(arr_n_[arr_n_ < 0].sum(), 3)
pn = round(arr_n_[arr_n_ > 0].sum(), 3)
an = round(abs(pn/ln), 3)
print('amount of loss no_ {0} amount of profits no_ {1} attitude no {2} amount of loss n_ '
'{3} amount of profits n_ {4} attitude n {5}'.format(lno, pno, ano, ln, pn, an))
print('arr_n_ > 0 amount {0} arr_n_ < 0 amount {1} arr_no_ > 0 amount {2} arr_no_ < 0 amount {3}'.
format((arr_n_ > 0).sum(), (arr_n_ < 0).sum(), (arr_no_ > 0).sum(), (arr_no_ < 0).sum()))

You can also calculate A / B - test. Which shows that there is no significant difference.

z_csore : 0.29180281268336133 p value : 0.3852186972909527
amount of loss no_ -0.49 amount of profits no_ 0.77 attitude no 1.571 amount of loss n_ -0.597 amount of profits n_ 1.096 attitude n 1.836
arr_n_ > 0 amount 162 arr_n_ < 0 amount 100 arr_no_ > 0 amount 169 arr_no_ < 0 amount 99

But, if we calculate the amount of losses, the amount of profits and their ratio, then we see There is a difference and it is noticeable. Which is further confirmed by the best Sharp Ratio and coefficient Pearson correlations. Based on this, I assume that the model with additional data chose more those labels where there were more profits, even taking into account the fact that its precision is slightly less than other models. The search for the best parameters was carried out by coefficient Pearson correlations.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Currency-Forecasting

Cumulative balance charts When forecasting in finance, derivatives of the price itself are usually used. Sometimes this has some effect, the improvements are barely noticeable, if not accidental. Famous expression: 'garbage in the input garbage out. You can try to look for signs not in the price itself. Central banks, various bureaus produce numerous macroeconomic indicators. I used them as additional features for forecasting currencies.

For each currency area, more than 700 types of data were obtained (for each currency a couple of about 1500). I tried to reduce the number of features using feature importance based on a random forest, the importance of permutations, etc. This did not give the desired effect. Probably because there is little information in the data. Then I started to add one feature to the main dataset and see if there are any improvements in the balance curve. So Thus, the number of signs was reduced tenfold. This job took a lot time. Then i began to cycle through these selected features adding three to the main dataset or applied a genetic algorithm to find the best combination. A variant with a genetic algorithm found the best combinations with a large number of features and over time it turned out that it was retraining (overfitting). In the above chart, the result obtained for EURUSD since 1991 (only sales).Sell and hold in orange. On the left, the model trained without adding macroeconomic indicators ('Depozit_no'), on the right with adding ('Depozit_n'). The model was trained in January 2020 and the settings have not changed since then. Classifier was used for classification: GradientBoostingClassifier. Signs news 1, 2, 3 do not plan to disclose.

To reproduce the code, the following libraries are needed: numpy, pandas, scipy, sklearn, matplotlib, seaborn, talib. You also need an include file: qnt.

By code:

1.The dataset.csv file is read, in which EURUSD daily quotes and three selected macroeconomic indicators (these are their momentums shifted by the required amount and split-normalized). Example: if the data is dated October 1st and becomes available on November 15th, then it is shifted by about 45 days.

2.The price series is converted into a series of technical analysis indicators. This data+ macroeconomic indicators create a training dataset.

3.Based on the classifier labels obtained, balance curves are built in dataframes df_n, df_no.

4.Various indicators, a graph of balances and distributions of transactions are displayed.

this is Depozit_no --- Pearson corr train 0.992 Pearson corr test 0.918
Sharp_Ratio_train 3.45 Sharp_Ratio_test 2.54 amount of deals 960
Train precision recall f1-score support
0 0.61 0.25 0.35 2807
1 0.54 0.84 0.66 2895
accuracy 0.55 5702
macro avg 0.57 0.55 0.51 5702
weighted avg 0.57 0.55 0.51 5702
Test precision recall f1-score support
0 0.56 0.17 0.27 1355
1 0.52 0.87 0.65 1383
accuracy 0.52 2738
macro avg 0.54 0.52 0.46 2738
weighted avg 0.54 0.52 0.46 2738
this is Depozit_n --- Pearson corr train 0.994 Pearson corr test 0.981
Sharp_Ratio_train 6.15 Sharp_Ratio_test 3.27 amount of deals 1027
Train precision recall f1-score support
0 0.73 0.30 0.43 2807
1 0.57 0.89 0.70 2895
accuracy 0.60 5702
macro avg 0.65 0.60 0.56 5702
weighted avg 0.65 0.60 0.56 5702
Test precision recall f1-score support
0 0.52 0.44 0.48 1355
1 0.53 0.61 0.56 1383
accuracy 0.52 2738
macro avg 0.52 0.52 0.52 2738
weighted avg 0.52 0.52 0.52 2738

distribution of deals

import scipy.stats
arr_no_ = arr_no[il[0]:]
arr_n_ = arr_n[il[1]:]
def ep(N, n):
p = n/N
sigma = math.sqrt(p * (1 - p)/N)
return p, sigma
def a_b_statistic(N_A, n_a, N_B, n_b):
p_A, sigma_A = ep(N_A, n_a)
p_B, sigma_B = ep(N_B, n_b)
return (p_B - p_A)/math.sqrt(sigma_A ** 2 + sigma_B ** 2)
z_csore = a_b_statistic(len(arr_n_), (arr_n_ > 0).sum(), len(arr_no_), (arr_no_ > 0).sum())
p_value = scipy.stats.norm.sf(abs(z_csore))
print('z_csore : ' + str(z_csore), 'p value : ' + str(p_value))
lno = round(arr_no_[arr_no_ < 0].sum(), 3)#amount of loss
pno = round(arr_no_[arr_no_ > 0].sum(), 3)#amount of profits
ano = round(abs(pno/lno), 3)#the ratio of the amount of profits to the amount of losses
ln = round(arr_n_[arr_n_ < 0].sum(), 3)
pn = round(arr_n_[arr_n_ > 0].sum(), 3)
an = round(abs(pn/ln), 3)
print('amount of loss no_ {0} amount of profits no_ {1} attitude no {2} amount of loss n_ '
'{3} amount of profits n_ {4} attitude n {5}'.format(lno, pno, ano, ln, pn, an))
print('arr_n_ > 0 amount {0} arr_n_ < 0 amount {1} arr_no_ > 0 amount {2} arr_no_ < 0 amount {3}'.
format((arr_n_ > 0).sum(), (arr_n_ < 0).sum(), (arr_no_ > 0).sum(), (arr_no_ < 0).sum()))

You can also calculate A / B - test. Which shows that there is no significant difference.

z_csore : 0.29180281268336133 p value : 0.3852186972909527
amount of loss no_ -0.49 amount of profits no_ 0.77 attitude no 1.571 amount of loss n_ -0.597 amount of profits n_ 1.096 attitude n 1.836
arr_n_ > 0 amount 162 arr_n_ < 0 amount 100 arr_no_ > 0 amount 169 arr_no_ < 0 amount 99

But, if we calculate the amount of losses, the amount of profits and their ratio, then we see There is a difference and it is noticeable. Which is further confirmed by the best Sharp Ratio and coefficient Pearson correlations. Based on this, I assume that the model with additional data chose more those labels where there were more profits, even taking into account the fact that its precision is slightly less than other models. The search for the best parameters was carried out by coefficient Pearson correlations.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Currency-Forecasting

Cumulative balance charts When forecasting in finance, derivatives of the price itself are usually used. Sometimes this has some effect, the improvements are barely noticeable, if not accidental. Famous expression: 'garbage in the input garbage out. You can try to look for signs not in the price itself. Central banks, various bureaus produce numerous macroeconomic indicators. I used them as additional features for forecasting currencies.

For each currency area, more than 700 types of data were obtained (for each currency a couple of about 1500). I tried to reduce the number of features using feature importance based on a random forest, the importance of permutations, etc. This did not give the desired effect. Probably because there is little information in the data. Then I started to add one feature to the main dataset and see if there are any improvements in the balance curve. So Thus, the number of signs was reduced tenfold. This job took a lot time. Then i began to cycle through these selected features adding three to the main dataset or applied a genetic algorithm to find the best combination. A variant with a genetic algorithm found the best combinations with a large number of features and over time it turned out that it was retraining (overfitting). In the above chart, the result obtained for EURUSD since 1991 (only sales).Sell and hold in orange. On the left, the model trained without adding macroeconomic indicators ('Depozit_no'), on the right with adding ('Depozit_n'). The model was trained in January 2020 and the settings have not changed since then. Classifier was used for classification: GradientBoostingClassifier. Signs news 1, 2, 3 do not plan to disclose.

To reproduce the code, the following libraries are needed: numpy, pandas, scipy, sklearn, matplotlib, seaborn, talib. You also need an include file: qnt.

By code:

1.The dataset.csv file is read, in which EURUSD daily quotes and three selected macroeconomic indicators (these are their momentums shifted by the required amount and split-normalized). Example: if the data is dated October 1st and becomes available on November 15th, then it is shifted by about 45 days.

2.The price series is converted into a series of technical analysis indicators. This data+ macroeconomic indicators create a training dataset.

3.Based on the classifier labels obtained, balance curves are built in dataframes df_n, df_no.

4.Various indicators, a graph of balances and distributions of transactions are displayed.

this is Depozit_no --- Pearson corr train 0.992 Pearson corr test 0.918
Sharp_Ratio_train 3.45 Sharp_Ratio_test 2.54 amount of deals 960
Train precision recall f1-score support
0 0.61 0.25 0.35 2807
1 0.54 0.84 0.66 2895
accuracy 0.55 5702
macro avg 0.57 0.55 0.51 5702
weighted avg 0.57 0.55 0.51 5702
Test precision recall f1-score support
0 0.56 0.17 0.27 1355
1 0.52 0.87 0.65 1383
accuracy 0.52 2738
macro avg 0.54 0.52 0.46 2738
weighted avg 0.54 0.52 0.46 2738
this is Depozit_n --- Pearson corr train 0.994 Pearson corr test 0.981
Sharp_Ratio_train 6.15 Sharp_Ratio_test 3.27 amount of deals 1027
Train precision recall f1-score support
0 0.73 0.30 0.43 2807
1 0.57 0.89 0.70 2895
accuracy 0.60 5702
macro avg 0.65 0.60 0.56 5702
weighted avg 0.65 0.60 0.56 5702
Test precision recall f1-score support
0 0.52 0.44 0.48 1355
1 0.53 0.61 0.56 1383
accuracy 0.52 2738
macro avg 0.52 0.52 0.52 2738
weighted avg 0.52 0.52 0.52 2738

distribution of deals

import scipy.stats
arr_no_ = arr_no[il[0]:]
arr_n_ = arr_n[il[1]:]
def ep(N, n):
p = n/N
sigma = math.sqrt(p * (1 - p)/N)
return p, sigma
def a_b_statistic(N_A, n_a, N_B, n_b):
p_A, sigma_A = ep(N_A, n_a)
p_B, sigma_B = ep(N_B, n_b)
return (p_B - p_A)/math.sqrt(sigma_A ** 2 + sigma_B ** 2)
z_csore = a_b_statistic(len(arr_n_), (arr_n_ > 0).sum(), len(arr_no_), (arr_no_ > 0).sum())
p_value = scipy.stats.norm.sf(abs(z_csore))
print('z_csore : ' + str(z_csore), 'p value : ' + str(p_value))
lno = round(arr_no_[arr_no_ < 0].sum(), 3)#amount of loss
pno = round(arr_no_[arr_no_ > 0].sum(), 3)#amount of profits
ano = round(abs(pno/lno), 3)#the ratio of the amount of profits to the amount of losses
ln = round(arr_n_[arr_n_ < 0].sum(), 3)
pn = round(arr_n_[arr_n_ > 0].sum(), 3)
an = round(abs(pn/ln), 3)
print('amount of loss no_ {0} amount of profits no_ {1} attitude no {2} amount of loss n_ '
'{3} amount of profits n_ {4} attitude n {5}'.format(lno, pno, ano, ln, pn, an))
print('arr_n_ > 0 amount {0} arr_n_ < 0 amount {1} arr_no_ > 0 amount {2} arr_no_ < 0 amount {3}'.
format((arr_n_ > 0).sum(), (arr_n_ < 0).sum(), (arr_no_ > 0).sum(), (arr_no_ < 0).sum()))

You can also calculate A / B - test. Which shows that there is no significant difference.

z_csore : 0.29180281268336133 p value : 0.3852186972909527
amount of loss no_ -0.49 amount of profits no_ 0.77 attitude no 1.571 amount of loss n_ -0.597 amount of profits n_ 1.096 attitude n 1.836
arr_n_ > 0 amount 162 arr_n_ < 0 amount 100 arr_no_ > 0 amount 169 arr_no_ < 0 amount 99

But, if we calculate the amount of losses, the amount of profits and their ratio, then we see There is a difference and it is noticeable. Which is further confirmed by the best Sharp Ratio and coefficient Pearson correlations. Based on this, I assume that the model with additional data chose more those labels where there were more profits, even taking into account the fact that its precision is slightly less than other models. The search for the best parameters was carried out by coefficient Pearson correlations.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Currency-Forecasting

Cumulative balance charts When forecasting in finance, derivatives of the price itself are usually used. Sometimes this has some effect, the improvements are barely noticeable, if not accidental. Famous expression: 'garbage in the input garbage out. You can try to look for signs not in the price itself. Central banks, various bureaus produce numerous macroeconomic indicators. I used them as additional features for forecasting currencies.

For each currency area, more than 700 types of data were obtained (for each currency a couple of about 1500). I tried to reduce the number of features using feature importance based on a random forest, the importance of permutations, etc. This did not give the desired effect. Probably because there is little information in the data. Then I started to add one feature to the main dataset and see if there are any improvements in the balance curve. So Thus, the number of signs was reduced tenfold. This job took a lot time. Then i began to cycle through these selected features adding three to the main dataset or applied a genetic algorithm to find the best combination. A variant with a genetic algorithm found the best combinations with a large number of features and over time it turned out that it was retraining (overfitting). In the above chart, the result obtained for EURUSD since 1991 (only sales).Sell and hold in orange. On the left, the model trained without adding macroeconomic indicators ('Depozit_no'), on the right with adding ('Depozit_n'). The model was trained in January 2020 and the settings have not changed since then. Classifier was used for classification: GradientBoostingClassifier. Signs news 1, 2, 3 do not plan to disclose.

To reproduce the code, the following libraries are needed: numpy, pandas, scipy, sklearn, matplotlib, seaborn, talib. You also need an include file: qnt.

By code:

1.The dataset.csv file is read, in which EURUSD daily quotes and three selected macroeconomic indicators (these are their momentums shifted by the required amount and split-normalized). Example: if the data is dated October 1st and becomes available on November 15th, then it is shifted by about 45 days.

2.The price series is converted into a series of technical analysis indicators. This data+ macroeconomic indicators create a training dataset.

3.Based on the classifier labels obtained, balance curves are built in dataframes df_n, df_no.

4.Various indicators, a graph of balances and distributions of transactions are displayed.

this is Depozit_no --- Pearson corr train 0.992 Pearson corr test 0.918
Sharp_Ratio_train 3.45 Sharp_Ratio_test 2.54 amount of deals 960
Train precision recall f1-score support
0 0.61 0.25 0.35 2807
1 0.54 0.84 0.66 2895
accuracy 0.55 5702
macro avg 0.57 0.55 0.51 5702
weighted avg 0.57 0.55 0.51 5702
Test precision recall f1-score support
0 0.56 0.17 0.27 1355
1 0.52 0.87 0.65 1383
accuracy 0.52 2738
macro avg 0.54 0.52 0.46 2738
weighted avg 0.54 0.52 0.46 2738
this is Depozit_n --- Pearson corr train 0.994 Pearson corr test 0.981
Sharp_Ratio_train 6.15 Sharp_Ratio_test 3.27 amount of deals 1027
Train precision recall f1-score support
0 0.73 0.30 0.43 2807
1 0.57 0.89 0.70 2895
accuracy 0.60 5702
macro avg 0.65 0.60 0.56 5702
weighted avg 0.65 0.60 0.56 5702
Test precision recall f1-score support
0 0.52 0.44 0.48 1355
1 0.53 0.61 0.56 1383
accuracy 0.52 2738
macro avg 0.52 0.52 0.52 2738
weighted avg 0.52 0.52 0.52 2738

distribution of deals

import scipy.stats
arr_no_ = arr_no[il[0]:]
arr_n_ = arr_n[il[1]:]
def ep(N, n):
p = n/N
sigma = math.sqrt(p * (1 - p)/N)
return p, sigma
def a_b_statistic(N_A, n_a, N_B, n_b):
p_A, sigma_A = ep(N_A, n_a)
p_B, sigma_B = ep(N_B, n_b)
return (p_B - p_A)/math.sqrt(sigma_A ** 2 + sigma_B ** 2)
z_csore = a_b_statistic(len(arr_n_), (arr_n_ > 0).sum(), len(arr_no_), (arr_no_ > 0).sum())
p_value = scipy.stats.norm.sf(abs(z_csore))
print('z_csore : ' + str(z_csore), 'p value : ' + str(p_value))
lno = round(arr_no_[arr_no_ < 0].sum(), 3)#amount of loss
pno = round(arr_no_[arr_no_ > 0].sum(), 3)#amount of profits
ano = round(abs(pno/lno), 3)#the ratio of the amount of profits to the amount of losses
ln = round(arr_n_[arr_n_ < 0].sum(), 3)
pn = round(arr_n_[arr_n_ > 0].sum(), 3)
an = round(abs(pn/ln), 3)
print('amount of loss no_ {0} amount of profits no_ {1} attitude no {2} amount of loss n_ '
'{3} amount of profits n_ {4} attitude n {5}'.format(lno, pno, ano, ln, pn, an))
print('arr_n_ > 0 amount {0} arr_n_ < 0 amount {1} arr_no_ > 0 amount {2} arr_no_ < 0 amount {3}'.
format((arr_n_ > 0).sum(), (arr_n_ < 0).sum(), (arr_no_ > 0).sum(), (arr_no_ < 0).sum()))

You can also calculate A / B - test. Which shows that there is no significant difference.

z_csore : 0.29180281268336133 p value : 0.3852186972909527
amount of loss no_ -0.49 amount of profits no_ 0.77 attitude no 1.571 amount of loss n_ -0.597 amount of profits n_ 1.096 attitude n 1.836
arr_n_ > 0 amount 162 arr_n_ < 0 amount 100 arr_no_ > 0 amount 169 arr_no_ < 0 amount 99

But, if we calculate the amount of losses, the amount of profits and their ratio, then we see There is a difference and it is noticeable. Which is further confirmed by the best Sharp Ratio and coefficient Pearson correlations. Based on this, I assume that the model with additional data chose more those labels where there were more profits, even taking into account the fact that its precision is slightly less than other models. The search for the best parameters was carried out by coefficient Pearson correlations.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Currency-Forecasting

Cumulative balance charts When forecasting in finance, derivatives of the price itself are usually used. Sometimes this has some effect, the improvements are barely noticeable, if not accidental. Famous expression: 'garbage in the input garbage out. You can try to look for signs not in the price itself. Central banks, various bureaus produce numerous macroeconomic indicators. I used them as additional features for forecasting currencies.

For each currency area, more than 700 types of data were obtained (for each currency a couple of about 1500). I tried to reduce the number of features using feature importance based on a random forest, the importance of permutations, etc. This did not give the desired effect. Probably because there is little information in the data. Then I started to add one feature to the main dataset and see if there are any improvements in the balance curve. So Thus, the number of signs was reduced tenfold. This job took a lot time. Then i began to cycle through these selected features adding three to the main dataset or applied a genetic algorithm to find the best combination. A variant with a genetic algorithm found the best combinations with a large number of features and over time it turned out that it was retraining (overfitting). In the above chart, the result obtained for EURUSD since 1991 (only sales).Sell and hold in orange. On the left, the model trained without adding macroeconomic indicators ('Depozit_no'), on the right with adding ('Depozit_n'). The model was trained in January 2020 and the settings have not changed since then. Classifier was used for classification: GradientBoostingClassifier. Signs news 1, 2, 3 do not plan to disclose.

To reproduce the code, the following libraries are needed: numpy, pandas, scipy, sklearn, matplotlib, seaborn, talib. You also need an include file: qnt.

By code:

1.The dataset.csv file is read, in which EURUSD daily quotes and three selected macroeconomic indicators (these are their momentums shifted by the required amount and split-normalized). Example: if the data is dated October 1st and becomes available on November 15th, then it is shifted by about 45 days.

2.The price series is converted into a series of technical analysis indicators. This data+ macroeconomic indicators create a training dataset.

3.Based on the classifier labels obtained, balance curves are built in dataframes df_n, df_no.

4.Various indicators, a graph of balances and distributions of transactions are displayed.

this is Depozit_no --- Pearson corr train 0.992 Pearson corr test 0.918
Sharp_Ratio_train 3.45 Sharp_Ratio_test 2.54 amount of deals 960
Train precision recall f1-score support
0 0.61 0.25 0.35 2807
1 0.54 0.84 0.66 2895
accuracy 0.55 5702
macro avg 0.57 0.55 0.51 5702
weighted avg 0.57 0.55 0.51 5702
Test precision recall f1-score support
0 0.56 0.17 0.27 1355
1 0.52 0.87 0.65 1383
accuracy 0.52 2738
macro avg 0.54 0.52 0.46 2738
weighted avg 0.54 0.52 0.46 2738
this is Depozit_n --- Pearson corr train 0.994 Pearson corr test 0.981
Sharp_Ratio_train 6.15 Sharp_Ratio_test 3.27 amount of deals 1027
Train precision recall f1-score support
0 0.73 0.30 0.43 2807
1 0.57 0.89 0.70 2895
accuracy 0.60 5702
macro avg 0.65 0.60 0.56 5702
weighted avg 0.65 0.60 0.56 5702
Test precision recall f1-score support
0 0.52 0.44 0.48 1355
1 0.53 0.61 0.56 1383
accuracy 0.52 2738
macro avg 0.52 0.52 0.52 2738
weighted avg 0.52 0.52 0.52 2738

distribution of deals

import scipy.stats
arr_no_ = arr_no[il[0]:]
arr_n_ = arr_n[il[1]:]
def ep(N, n):
p = n/N
sigma = math.sqrt(p * (1 - p)/N)
return p, sigma
def a_b_statistic(N_A, n_a, N_B, n_b):
p_A, sigma_A = ep(N_A, n_a)
p_B, sigma_B = ep(N_B, n_b)
return (p_B - p_A)/math.sqrt(sigma_A ** 2 + sigma_B ** 2)
z_csore = a_b_statistic(len(arr_n_), (arr_n_ > 0).sum(), len(arr_no_), (arr_no_ > 0).sum())
p_value = scipy.stats.norm.sf(abs(z_csore))
print('z_csore : ' + str(z_csore), 'p value : ' + str(p_value))
lno = round(arr_no_[arr_no_ < 0].sum(), 3)#amount of loss
pno = round(arr_no_[arr_no_ > 0].sum(), 3)#amount of profits
ano = round(abs(pno/lno), 3)#the ratio of the amount of profits to the amount of losses
ln = round(arr_n_[arr_n_ < 0].sum(), 3)
pn = round(arr_n_[arr_n_ > 0].sum(), 3)
an = round(abs(pn/ln), 3)
print('amount of loss no_ {0} amount of profits no_ {1} attitude no {2} amount of loss n_ '
'{3} amount of profits n_ {4} attitude n {5}'.format(lno, pno, ano, ln, pn, an))
print('arr_n_ > 0 amount {0} arr_n_ < 0 amount {1} arr_no_ > 0 amount {2} arr_no_ < 0 amount {3}'.
format((arr_n_ > 0).sum(), (arr_n_ < 0).sum(), (arr_no_ > 0).sum(), (arr_no_ < 0).sum()))

You can also calculate A / B - test. Which shows that there is no significant difference.

z_csore : 0.29180281268336133 p value : 0.3852186972909527
amount of loss no_ -0.49 amount of profits no_ 0.77 attitude no 1.571 amount of loss n_ -0.597 amount of profits n_ 1.096 attitude n 1.836
arr_n_ > 0 amount 162 arr_n_ < 0 amount 100 arr_no_ > 0 amount 169 arr_no_ < 0 amount 99

But, if we calculate the amount of losses, the amount of profits and their ratio, then we see There is a difference and it is noticeable. Which is further confirmed by the best Sharp Ratio and coefficient Pearson correlations. Based on this, I assume that the model with additional data chose more those labels where there were more profits, even taking into account the fact that its precision is slightly less than other models. The search for the best parameters was carried out by coefficient Pearson correlations.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Currency-Forecasting

Cumulative balance charts When forecasting in finance, derivatives of the price itself are usually used. Sometimes this has some effect, the improvements are barely noticeable, if not accidental. Famous expression: 'garbage in the input garbage out. You can try to look for signs not in the price itself. Central banks, various bureaus produce numerous macroeconomic indicators. I used them as additional features for forecasting currencies.

For each currency area, more than 700 types of data were obtained (for each currency a couple of about 1500). I tried to reduce the number of features using feature importance based on a random forest, the importance of permutations, etc. This did not give the desired effect. Probably because there is little information in the data. Then I started to add one feature to the main dataset and see if there are any improvements in the balance curve. So Thus, the number of signs was reduced tenfold. This job took a lot time. Then i began to cycle through these selected features adding three to the main dataset or applied a genetic algorithm to find the best combination. A variant with a genetic algorithm found the best combinations with a large number of features and over time it turned out that it was retraining (overfitting). In the above chart, the result obtained for EURUSD since 1991 (only sales).Sell and hold in orange. On the left, the model trained without adding macroeconomic indicators ('Depozit_no'), on the right with adding ('Depozit_n'). The model was trained in January 2020 and the settings have not changed since then. Classifier was used for classification: GradientBoostingClassifier. Signs news 1, 2, 3 do not plan to disclose.

To reproduce the code, the following libraries are needed: numpy, pandas, scipy, sklearn, matplotlib, seaborn, talib. You also need an include file: qnt.

By code:

1.The dataset.csv file is read, in which EURUSD daily quotes and three selected macroeconomic indicators (these are their momentums shifted by the required amount and split-normalized). Example: if the data is dated October 1st and becomes available on November 15th, then it is shifted by about 45 days.

2.The price series is converted into a series of technical analysis indicators. This data+ macroeconomic indicators create a training dataset.

3.Based on the classifier labels obtained, balance curves are built in dataframes df_n, df_no.

4.Various indicators, a graph of balances and distributions of transactions are displayed.

this is Depozit_no --- Pearson corr train 0.992 Pearson corr test 0.918
Sharp_Ratio_train 3.45 Sharp_Ratio_test 2.54 amount of deals 960
Train precision recall f1-score support
0 0.61 0.25 0.35 2807
1 0.54 0.84 0.66 2895
accuracy 0.55 5702
macro avg 0.57 0.55 0.51 5702
weighted avg 0.57 0.55 0.51 5702
Test precision recall f1-score support
0 0.56 0.17 0.27 1355
1 0.52 0.87 0.65 1383
accuracy 0.52 2738
macro avg 0.54 0.52 0.46 2738
weighted avg 0.54 0.52 0.46 2738
this is Depozit_n --- Pearson corr train 0.994 Pearson corr test 0.981
Sharp_Ratio_train 6.15 Sharp_Ratio_test 3.27 amount of deals 1027
Train precision recall f1-score support
0 0.73 0.30 0.43 2807
1 0.57 0.89 0.70 2895
accuracy 0.60 5702
macro avg 0.65 0.60 0.56 5702
weighted avg 0.65 0.60 0.56 5702
Test precision recall f1-score support
0 0.52 0.44 0.48 1355
1 0.53 0.61 0.56 1383
accuracy 0.52 2738
macro avg 0.52 0.52 0.52 2738
weighted avg 0.52 0.52 0.52 2738

distribution of deals

import scipy.stats
arr_no_ = arr_no[il[0]:]
arr_n_ = arr_n[il[1]:]
def ep(N, n):
p = n/N
sigma = math.sqrt(p * (1 - p)/N)
return p, sigma
def a_b_statistic(N_A, n_a, N_B, n_b):
p_A, sigma_A = ep(N_A, n_a)
p_B, sigma_B = ep(N_B, n_b)
return (p_B - p_A)/math.sqrt(sigma_A ** 2 + sigma_B ** 2)
z_csore = a_b_statistic(len(arr_n_), (arr_n_ > 0).sum(), len(arr_no_), (arr_no_ > 0).sum())
p_value = scipy.stats.norm.sf(abs(z_csore))
print('z_csore : ' + str(z_csore), 'p value : ' + str(p_value))
lno = round(arr_no_[arr_no_ < 0].sum(), 3)#amount of loss
pno = round(arr_no_[arr_no_ > 0].sum(), 3)#amount of profits
ano = round(abs(pno/lno), 3)#the ratio of the amount of profits to the amount of losses
ln = round(arr_n_[arr_n_ < 0].sum(), 3)
pn = round(arr_n_[arr_n_ > 0].sum(), 3)
an = round(abs(pn/ln), 3)
print('amount of loss no_ {0} amount of profits no_ {1} attitude no {2} amount of loss n_ '
'{3} amount of profits n_ {4} attitude n {5}'.format(lno, pno, ano, ln, pn, an))
print('arr_n_ > 0 amount {0} arr_n_ < 0 amount {1} arr_no_ > 0 amount {2} arr_no_ < 0 amount {3}'.
format((arr_n_ > 0).sum(), (arr_n_ < 0).sum(), (arr_no_ > 0).sum(), (arr_no_ < 0).sum()))

You can also calculate A / B - test. Which shows that there is no significant difference.

z_csore : 0.29180281268336133 p value : 0.3852186972909527
amount of loss no_ -0.49 amount of profits no_ 0.77 attitude no 1.571 amount of loss n_ -0.597 amount of profits n_ 1.096 attitude n 1.836
arr_n_ > 0 amount 162 arr_n_ < 0 amount 100 arr_no_ > 0 amount 169 arr_no_ < 0 amount 99

But, if we calculate the amount of losses, the amount of profits and their ratio, then we see There is a difference and it is noticeable. Which is further confirmed by the best Sharp Ratio and coefficient Pearson correlations. Based on this, I assume that the model with additional data chose more those labels where there were more profits, even taking into account the fact that its precision is slightly less than other models. The search for the best parameters was carried out by coefficient Pearson correlations.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Currency-Forecasting

Cumulative balance charts When forecasting in finance, derivatives of the price itself are usually used. Sometimes this has some effect, the improvements are barely noticeable, if not accidental. Famous expression: 'garbage in the input garbage out. You can try to look for signs not in the price itself. Central banks, various bureaus produce numerous macroeconomic indicators. I used them as additional features for forecasting currencies.

For each currency area, more than 700 types of data were obtained (for each currency a couple of about 1500). I tried to reduce the number of features using feature importance based on a random forest, the importance of permutations, etc. This did not give the desired effect. Probably because there is little information in the data. Then I started to add one feature to the main dataset and see if there are any improvements in the balance curve. So Thus, the number of signs was reduced tenfold. This job took a lot time. Then i began to cycle through these selected features adding three to the main dataset or applied a genetic algorithm to find the best combination. A variant with a genetic algorithm found the best combinations with a large number of features and over time it turned out that it was retraining (overfitting). In the above chart, the result obtained for EURUSD since 1991 (only sales).Sell and hold in orange. On the left, the model trained without adding macroeconomic indicators ('Depozit_no'), on the right with adding ('Depozit_n'). The model was trained in January 2020 and the settings have not changed since then. Classifier was used for classification: GradientBoostingClassifier. Signs news 1, 2, 3 do not plan to disclose.

To reproduce the code, the following libraries are needed: numpy, pandas, scipy, sklearn, matplotlib, seaborn, talib. You also need an include file: qnt.

By code:

1.The dataset.csv file is read, in which EURUSD daily quotes and three selected macroeconomic indicators (these are their momentums shifted by the required amount and split-normalized). Example: if the data is dated October 1st and becomes available on November 15th, then it is shifted by about 45 days.

2.The price series is converted into a series of technical analysis indicators. This data+ macroeconomic indicators create a training dataset.

3.Based on the classifier labels obtained, balance curves are built in dataframes df_n, df_no.

4.Various indicators, a graph of balances and distributions of transactions are displayed.

this is Depozit_no --- Pearson corr train 0.992 Pearson corr test 0.918
Sharp_Ratio_train 3.45 Sharp_Ratio_test 2.54 amount of deals 960
Train precision recall f1-score support
0 0.61 0.25 0.35 2807
1 0.54 0.84 0.66 2895
accuracy 0.55 5702
macro avg 0.57 0.55 0.51 5702
weighted avg 0.57 0.55 0.51 5702
Test precision recall f1-score support
0 0.56 0.17 0.27 1355
1 0.52 0.87 0.65 1383
accuracy 0.52 2738
macro avg 0.54 0.52 0.46 2738
weighted avg 0.54 0.52 0.46 2738
this is Depozit_n --- Pearson corr train 0.994 Pearson corr test 0.981
Sharp_Ratio_train 6.15 Sharp_Ratio_test 3.27 amount of deals 1027
Train precision recall f1-score support
0 0.73 0.30 0.43 2807
1 0.57 0.89 0.70 2895
accuracy 0.60 5702
macro avg 0.65 0.60 0.56 5702
weighted avg 0.65 0.60 0.56 5702
Test precision recall f1-score support
0 0.52 0.44 0.48 1355
1 0.53 0.61 0.56 1383
accuracy 0.52 2738
macro avg 0.52 0.52 0.52 2738
weighted avg 0.52 0.52 0.52 2738

distribution of deals

import scipy.stats
arr_no_ = arr_no[il[0]:]
arr_n_ = arr_n[il[1]:]
def ep(N, n):
p = n/N
sigma = math.sqrt(p * (1 - p)/N)
return p, sigma
def a_b_statistic(N_A, n_a, N_B, n_b):
p_A, sigma_A = ep(N_A, n_a)
p_B, sigma_B = ep(N_B, n_b)
return (p_B - p_A)/math.sqrt(sigma_A ** 2 + sigma_B ** 2)
z_csore = a_b_statistic(len(arr_n_), (arr_n_ > 0).sum(), len(arr_no_), (arr_no_ > 0).sum())
p_value = scipy.stats.norm.sf(abs(z_csore))
print('z_csore : ' + str(z_csore), 'p value : ' + str(p_value))
lno = round(arr_no_[arr_no_ < 0].sum(), 3)#amount of loss
pno = round(arr_no_[arr_no_ > 0].sum(), 3)#amount of profits
ano = round(abs(pno/lno), 3)#the ratio of the amount of profits to the amount of losses
ln = round(arr_n_[arr_n_ < 0].sum(), 3)
pn = round(arr_n_[arr_n_ > 0].sum(), 3)
an = round(abs(pn/ln), 3)
print('amount of loss no_ {0} amount of profits no_ {1} attitude no {2} amount of loss n_ '
'{3} amount of profits n_ {4} attitude n {5}'.format(lno, pno, ano, ln, pn, an))
print('arr_n_ > 0 amount {0} arr_n_ < 0 amount {1} arr_no_ > 0 amount {2} arr_no_ < 0 amount {3}'.
format((arr_n_ > 0).sum(), (arr_n_ < 0).sum(), (arr_no_ > 0).sum(), (arr_no_ < 0).sum()))

You can also calculate A / B - test. Which shows that there is no significant difference.

z_csore : 0.29180281268336133 p value : 0.3852186972909527
amount of loss no_ -0.49 amount of profits no_ 0.77 attitude no 1.571 amount of loss n_ -0.597 amount of profits n_ 1.096 attitude n 1.836
arr_n_ > 0 amount 162 arr_n_ < 0 amount 100 arr_no_ > 0 amount 169 arr_no_ < 0 amount 99

But, if we calculate the amount of losses, the amount of profits and their ratio, then we see There is a difference and it is noticeable. Which is further confirmed by the best Sharp Ratio and coefficient Pearson correlations. Based on this, I assume that the model with additional data chose more those labels where there were more profits, even taking into account the fact that its precision is slightly less than other models. The search for the best parameters was carried out by coefficient Pearson correlations.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Currency-Forecasting

Cumulative balance charts When forecasting in finance, derivatives of the price itself are usually used. Sometimes this has some effect, the improvements are barely noticeable, if not accidental. Famous expression: 'garbage in the input garbage out. You can try to look for signs not in the price itself. Central banks, various bureaus produce numerous macroeconomic indicators. I used them as additional features for forecasting currencies.

For each currency area, more than 700 types of data were obtained (for each currency a couple of about 1500). I tried to reduce the number of features using feature importance based on a random forest, the importance of permutations, etc. This did not give the desired effect. Probably because there is little information in the data. Then I started to add one feature to the main dataset and see if there are any improvements in the balance curve. So Thus, the number of signs was reduced tenfold. This job took a lot time. Then i began to cycle through these selected features adding three to the main dataset or applied a genetic algorithm to find the best combination. A variant with a genetic algorithm found the best combinations with a large number of features and over time it turned out that it was retraining (overfitting). In the above chart, the result obtained for EURUSD since 1991 (only sales).Sell and hold in orange. On the left, the model trained without adding macroeconomic indicators ('Depozit_no'), on the right with adding ('Depozit_n'). The model was trained in January 2020 and the settings have not changed since then. Classifier was used for classification: GradientBoostingClassifier. Signs news 1, 2, 3 do not plan to disclose.

To reproduce the code, the following libraries are needed: numpy, pandas, scipy, sklearn, matplotlib, seaborn, talib. You also need an include file: qnt.

By code:

1.The dataset.csv file is read, in which EURUSD daily quotes and three selected macroeconomic indicators (these are their momentums shifted by the required amount and split-normalized). Example: if the data is dated October 1st and becomes available on November 15th, then it is shifted by about 45 days.

2.The price series is converted into a series of technical analysis indicators. This data+ macroeconomic indicators create a training dataset.

3.Based on the classifier labels obtained, balance curves are built in dataframes df_n, df_no.

4.Various indicators, a graph of balances and distributions of transactions are displayed.

this is Depozit_no --- Pearson corr train 0.992 Pearson corr test 0.918
Sharp_Ratio_train 3.45 Sharp_Ratio_test 2.54 amount of deals 960
Train precision recall f1-score support
0 0.61 0.25 0.35 2807
1 0.54 0.84 0.66 2895
accuracy 0.55 5702
macro avg 0.57 0.55 0.51 5702
weighted avg 0.57 0.55 0.51 5702
Test precision recall f1-score support
0 0.56 0.17 0.27 1355
1 0.52 0.87 0.65 1383
accuracy 0.52 2738
macro avg 0.54 0.52 0.46 2738
weighted avg 0.54 0.52 0.46 2738
this is Depozit_n --- Pearson corr train 0.994 Pearson corr test 0.981
Sharp_Ratio_train 6.15 Sharp_Ratio_test 3.27 amount of deals 1027
Train precision recall f1-score support
0 0.73 0.30 0.43 2807
1 0.57 0.89 0.70 2895
accuracy 0.60 5702
macro avg 0.65 0.60 0.56 5702
weighted avg 0.65 0.60 0.56 5702
Test precision recall f1-score support
0 0.52 0.44 0.48 1355
1 0.53 0.61 0.56 1383
accuracy 0.52 2738
macro avg 0.52 0.52 0.52 2738
weighted avg 0.52 0.52 0.52 2738

distribution of deals

import scipy.stats
arr_no_ = arr_no[il[0]:]
arr_n_ = arr_n[il[1]:]
def ep(N, n):
p = n/N
sigma = math.sqrt(p * (1 - p)/N)
return p, sigma
def a_b_statistic(N_A, n_a, N_B, n_b):
p_A, sigma_A = ep(N_A, n_a)
p_B, sigma_B = ep(N_B, n_b)
return (p_B - p_A)/math.sqrt(sigma_A ** 2 + sigma_B ** 2)
z_csore = a_b_statistic(len(arr_n_), (arr_n_ > 0).sum(), len(arr_no_), (arr_no_ > 0).sum())
p_value = scipy.stats.norm.sf(abs(z_csore))
print('z_csore : ' + str(z_csore), 'p value : ' + str(p_value))
lno = round(arr_no_[arr_no_ < 0].sum(), 3)#amount of loss
pno = round(arr_no_[arr_no_ > 0].sum(), 3)#amount of profits
ano = round(abs(pno/lno), 3)#the ratio of the amount of profits to the amount of losses
ln = round(arr_n_[arr_n_ < 0].sum(), 3)
pn = round(arr_n_[arr_n_ > 0].sum(), 3)
an = round(abs(pn/ln), 3)
print('amount of loss no_ {0} amount of profits no_ {1} attitude no {2} amount of loss n_ '
'{3} amount of profits n_ {4} attitude n {5}'.format(lno, pno, ano, ln, pn, an))
print('arr_n_ > 0 amount {0} arr_n_ < 0 amount {1} arr_no_ > 0 amount {2} arr_no_ < 0 amount {3}'.
format((arr_n_ > 0).sum(), (arr_n_ < 0).sum(), (arr_no_ > 0).sum(), (arr_no_ < 0).sum()))

You can also calculate A / B - test. Which shows that there is no significant difference.

z_csore : 0.29180281268336133 p value : 0.3852186972909527
amount of loss no_ -0.49 amount of profits no_ 0.77 attitude no 1.571 amount of loss n_ -0.597 amount of profits n_ 1.096 attitude n 1.836
arr_n_ > 0 amount 162 arr_n_ < 0 amount 100 arr_no_ > 0 amount 169 arr_no_ < 0 amount 99

But, if we calculate the amount of losses, the amount of profits and their ratio, then we see There is a difference and it is noticeable. Which is further confirmed by the best Sharp Ratio and coefficient Pearson correlations. Based on this, I assume that the model with additional data chose more those labels where there were more profits, even taking into account the fact that its precision is slightly less than other models. The search for the best parameters was carried out by coefficient Pearson correlations.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages