Skip to content

Repository files navigation

VirtualPetAI

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Features

  • (Modified) Reinforcement Learning (RL) with Q-Learning: The pet AI learns over time by maximizing future rewards based on user interactions. However, it's not really true Q-Learning because the user still decides the actions. The pet AI learns which actions get the best approval of the user.
  • Sentiment Analysis: The pet AI can interpret the tone of the text you type and changes its state based on it.
  • State System: The pet AI has different states like "Happy", "Sad", "Excited", etc., based on the emotionScore, which is affected by actions taken by the user towards the pet AI

How it works

1. Q-Learning

The Q-learning formula I'm using is:

$Q(s, a) = (1 - α) * Q(s, a) + α * (reward + γ * maxFutureReward)$

Where:

  • $Q(s, a)$ represents the expected future reward for taking action a in state s.
  • α (Alpha) is the learning rate, controlling how much new information influences the Q-value (0 < α ≤ 1).
    • If α is close to 0, the pet AI learns slowly, relying more on past experiences for behavior
    • If α is close to 1, the pet AI learns quickly, relying more on new experiences for behavior
  • reward is the immediate reward for taking action a in state s.
  • γ (Gamma) is the discount factor, determining how much future rewards matter.
    • If γ = 0, only immediate rewards are considered.
    • If γ is close to 1, future rewards are given more importance.
  • maxFutureReward is the best possible reward achievable from the next state.

To be honest, this took the most amount of time for me to wrap my head around since I'm not the best at math, and I still don't fully understand it.

This formula is primarily what decides how the pet AI will react to the user's actions (Feed, Praise, Scold, etc.)

2. Sentiment Analysis

In this project, I'm using ML.NET an open-source machine learning framework.

For my sentiment analysis model, I decided to use text classification, where the model receives an input text (ex: "Bad boy! Stop chewing that!") and is asked to guess the output ("Positive, "Neutral", "Negative").

In order to predict the sentiment of text, I had to somehow find a ton of labeled data to feed the model. I couldn't find much of anything for sentiment analysis data based on humans interacting with pets, so I used synthetic data (basically, I prompted ChatGPT to give me a bunch of labeled sentiment analysis data)

After that, in order to train the model, I had to pick the best classifier to determine which label best fits input given to the model. I went with the SdcaMaximumEntropy classifier, since it's effective for multi-class classification problems like sentiment analysis

Behavior when ignored

This is part of the wider state system, but the behavior when ignored is one of my favorite parts of the project so I think it should get it's own mini-section

If you pick the action "Ignore" three times in a row, the pet AI will recognize that it's being neglected and will perform an action to try to get attention from the user, like spinning around, jumping, or even whimpering.

It decides the action with an Epsilon-Greedy Algorithm, which is a common strategy in RL to balance between exploration (trying new actions) and exploitation (choosing the action with the highest expected reward). The pet AI has a 20% chance to explore (choosing a random action) and an 80% chance to exploit (choosing the action with the highest Q-value)

The Q-values reflect the expected reward for an action. The better the action is considered in the context of the pet's emotional state and its goal of getting the user's attention.

Example demonstration

Interact with your virtual AI pet! Type "Feed", "Play", "Praise", "Ignore", "Yell", "Take away toy", "Scold" or chat with it. Type "exit" to exit the program.
> Feed
Pet reacts to Feed: Reward 3
Pet state: Neutral | Emotion Score: 3
> Play
Pet reacts to Play: Reward 3.18
Pet state: Happy | Emotion Score: 6
> Praise
Pet reacts to Praise: Reward 5.6
Pet state: Happy | Emotion Score: 11
> You're a good doggo :)
Pet detects sentiment: Positive
Positive score: 88.2% | Neutral score: 0.9% | Negative score: 10.9%
Pet state: Happy | Emotion Score: 14

About

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all \x3Cpre>\x3Ccode> blocks (function() { function addCopyButtons() { document.querySelectorAll('pre code').forEach(function(codeBlock) { if (codeBlock.parentElement.hasAttribute('data-copy-added')) return; codeBlock.parentElement.setAttribute('data-copy-added', 'true'); var btn = document.createElement('button'); btn.textContent = 'Copy'; btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;'; btn.onmouseover = function() { this.style.opacity = '1'; }; btn.onmouseout = function() { this.style.opacity = '0.7'; }; btn.onclick = function() { navigator.clipboard.writeText(codeBlock.textContent).then(function() { btn.textContent = 'Copied!'; setTimeout(function() { btn.textContent = 'Copy'; }, 1500); }); }; codeBlock.parentElement.style.position = 'relative'; codeBlock.parentElement.appendChild(btn); }); } addCopyButtons(); // Re-run on dynamic content var observer = new MutationObserver(addCopyButtons); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + ' GitHub - rosalgie/VirtualPetAI: Pet AI using Q-Learning and Sentiment Analysis in ML.NET · GitHub
Skip to content

Repository files navigation

VirtualPetAI

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Features

  • (Modified) Reinforcement Learning (RL) with Q-Learning: The pet AI learns over time by maximizing future rewards based on user interactions. However, it's not really true Q-Learning because the user still decides the actions. The pet AI learns which actions get the best approval of the user.
  • Sentiment Analysis: The pet AI can interpret the tone of the text you type and changes its state based on it.
  • State System: The pet AI has different states like "Happy", "Sad", "Excited", etc., based on the emotionScore, which is affected by actions taken by the user towards the pet AI

How it works

1. Q-Learning

The Q-learning formula I'm using is:

$Q(s, a) = (1 - α) * Q(s, a) + α * (reward + γ * maxFutureReward)$

Where:

  • $Q(s, a)$ represents the expected future reward for taking action a in state s.
  • α (Alpha) is the learning rate, controlling how much new information influences the Q-value (0 < α ≤ 1).
    • If α is close to 0, the pet AI learns slowly, relying more on past experiences for behavior
    • If α is close to 1, the pet AI learns quickly, relying more on new experiences for behavior
  • reward is the immediate reward for taking action a in state s.
  • γ (Gamma) is the discount factor, determining how much future rewards matter.
    • If γ = 0, only immediate rewards are considered.
    • If γ is close to 1, future rewards are given more importance.
  • maxFutureReward is the best possible reward achievable from the next state.

To be honest, this took the most amount of time for me to wrap my head around since I'm not the best at math, and I still don't fully understand it.

This formula is primarily what decides how the pet AI will react to the user's actions (Feed, Praise, Scold, etc.)

2. Sentiment Analysis

In this project, I'm using ML.NET an open-source machine learning framework.

For my sentiment analysis model, I decided to use text classification, where the model receives an input text (ex: "Bad boy! Stop chewing that!") and is asked to guess the output ("Positive, "Neutral", "Negative").

In order to predict the sentiment of text, I had to somehow find a ton of labeled data to feed the model. I couldn't find much of anything for sentiment analysis data based on humans interacting with pets, so I used synthetic data (basically, I prompted ChatGPT to give me a bunch of labeled sentiment analysis data)

After that, in order to train the model, I had to pick the best classifier to determine which label best fits input given to the model. I went with the SdcaMaximumEntropy classifier, since it's effective for multi-class classification problems like sentiment analysis

Behavior when ignored

This is part of the wider state system, but the behavior when ignored is one of my favorite parts of the project so I think it should get it's own mini-section

If you pick the action "Ignore" three times in a row, the pet AI will recognize that it's being neglected and will perform an action to try to get attention from the user, like spinning around, jumping, or even whimpering.

It decides the action with an Epsilon-Greedy Algorithm, which is a common strategy in RL to balance between exploration (trying new actions) and exploitation (choosing the action with the highest expected reward). The pet AI has a 20% chance to explore (choosing a random action) and an 80% chance to exploit (choosing the action with the highest Q-value)

The Q-values reflect the expected reward for an action. The better the action is considered in the context of the pet's emotional state and its goal of getting the user's attention.

Example demonstration

Interact with your virtual AI pet! Type "Feed", "Play", "Praise", "Ignore", "Yell", "Take away toy", "Scold" or chat with it. Type "exit" to exit the program.
> Feed
Pet reacts to Feed: Reward 3
Pet state: Neutral | Emotion Score: 3
> Play
Pet reacts to Play: Reward 3.18
Pet state: Happy | Emotion Score: 6
> Praise
Pet reacts to Praise: Reward 5.6
Pet state: Happy | Emotion Score: 11
> You're a good doggo :)
Pet detects sentiment: Positive
Positive score: 88.2% | Neutral score: 0.9% | Negative score: 10.9%
Pet state: Happy | Emotion Score: 14

About

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - rosalgie/VirtualPetAI: Pet AI using Q-Learning and Sentiment Analysis in ML.NET · GitHub
Skip to content

Repository files navigation

VirtualPetAI

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Features

  • (Modified) Reinforcement Learning (RL) with Q-Learning: The pet AI learns over time by maximizing future rewards based on user interactions. However, it's not really true Q-Learning because the user still decides the actions. The pet AI learns which actions get the best approval of the user.
  • Sentiment Analysis: The pet AI can interpret the tone of the text you type and changes its state based on it.
  • State System: The pet AI has different states like "Happy", "Sad", "Excited", etc., based on the emotionScore, which is affected by actions taken by the user towards the pet AI

How it works

1. Q-Learning

The Q-learning formula I'm using is:

$Q(s, a) = (1 - α) * Q(s, a) + α * (reward + γ * maxFutureReward)$

Where:

  • $Q(s, a)$ represents the expected future reward for taking action a in state s.
  • α (Alpha) is the learning rate, controlling how much new information influences the Q-value (0 < α ≤ 1).
    • If α is close to 0, the pet AI learns slowly, relying more on past experiences for behavior
    • If α is close to 1, the pet AI learns quickly, relying more on new experiences for behavior
  • reward is the immediate reward for taking action a in state s.
  • γ (Gamma) is the discount factor, determining how much future rewards matter.
    • If γ = 0, only immediate rewards are considered.
    • If γ is close to 1, future rewards are given more importance.
  • maxFutureReward is the best possible reward achievable from the next state.

To be honest, this took the most amount of time for me to wrap my head around since I'm not the best at math, and I still don't fully understand it.

This formula is primarily what decides how the pet AI will react to the user's actions (Feed, Praise, Scold, etc.)

2. Sentiment Analysis

In this project, I'm using ML.NET an open-source machine learning framework.

For my sentiment analysis model, I decided to use text classification, where the model receives an input text (ex: "Bad boy! Stop chewing that!") and is asked to guess the output ("Positive, "Neutral", "Negative").

In order to predict the sentiment of text, I had to somehow find a ton of labeled data to feed the model. I couldn't find much of anything for sentiment analysis data based on humans interacting with pets, so I used synthetic data (basically, I prompted ChatGPT to give me a bunch of labeled sentiment analysis data)

After that, in order to train the model, I had to pick the best classifier to determine which label best fits input given to the model. I went with the SdcaMaximumEntropy classifier, since it's effective for multi-class classification problems like sentiment analysis

Behavior when ignored

This is part of the wider state system, but the behavior when ignored is one of my favorite parts of the project so I think it should get it's own mini-section

If you pick the action "Ignore" three times in a row, the pet AI will recognize that it's being neglected and will perform an action to try to get attention from the user, like spinning around, jumping, or even whimpering.

It decides the action with an Epsilon-Greedy Algorithm, which is a common strategy in RL to balance between exploration (trying new actions) and exploitation (choosing the action with the highest expected reward). The pet AI has a 20% chance to explore (choosing a random action) and an 80% chance to exploit (choosing the action with the highest Q-value)

The Q-values reflect the expected reward for an action. The better the action is considered in the context of the pet's emotional state and its goal of getting the user's attention.

Example demonstration

Interact with your virtual AI pet! Type "Feed", "Play", "Praise", "Ignore", "Yell", "Take away toy", "Scold" or chat with it. Type "exit" to exit the program.
> Feed
Pet reacts to Feed: Reward 3
Pet state: Neutral | Emotion Score: 3
> Play
Pet reacts to Play: Reward 3.18
Pet state: Happy | Emotion Score: 6
> Praise
Pet reacts to Praise: Reward 5.6
Pet state: Happy | Emotion Score: 11
> You're a good doggo :)
Pet detects sentiment: Positive
Positive score: 88.2% | Neutral score: 0.9% | Negative score: 10.9%
Pet state: Happy | Emotion Score: 14

About

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - rosalgie/VirtualPetAI: Pet AI using Q-Learning and Sentiment Analysis in ML.NET · GitHub
Skip to content

Repository files navigation

VirtualPetAI

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Features

  • (Modified) Reinforcement Learning (RL) with Q-Learning: The pet AI learns over time by maximizing future rewards based on user interactions. However, it's not really true Q-Learning because the user still decides the actions. The pet AI learns which actions get the best approval of the user.
  • Sentiment Analysis: The pet AI can interpret the tone of the text you type and changes its state based on it.
  • State System: The pet AI has different states like "Happy", "Sad", "Excited", etc., based on the emotionScore, which is affected by actions taken by the user towards the pet AI

How it works

1. Q-Learning

The Q-learning formula I'm using is:

$Q(s, a) = (1 - α) * Q(s, a) + α * (reward + γ * maxFutureReward)$

Where:

  • $Q(s, a)$ represents the expected future reward for taking action a in state s.
  • α (Alpha) is the learning rate, controlling how much new information influences the Q-value (0 < α ≤ 1).
    • If α is close to 0, the pet AI learns slowly, relying more on past experiences for behavior
    • If α is close to 1, the pet AI learns quickly, relying more on new experiences for behavior
  • reward is the immediate reward for taking action a in state s.
  • γ (Gamma) is the discount factor, determining how much future rewards matter.
    • If γ = 0, only immediate rewards are considered.
    • If γ is close to 1, future rewards are given more importance.
  • maxFutureReward is the best possible reward achievable from the next state.

To be honest, this took the most amount of time for me to wrap my head around since I'm not the best at math, and I still don't fully understand it.

This formula is primarily what decides how the pet AI will react to the user's actions (Feed, Praise, Scold, etc.)

2. Sentiment Analysis

In this project, I'm using ML.NET an open-source machine learning framework.

For my sentiment analysis model, I decided to use text classification, where the model receives an input text (ex: "Bad boy! Stop chewing that!") and is asked to guess the output ("Positive, "Neutral", "Negative").

In order to predict the sentiment of text, I had to somehow find a ton of labeled data to feed the model. I couldn't find much of anything for sentiment analysis data based on humans interacting with pets, so I used synthetic data (basically, I prompted ChatGPT to give me a bunch of labeled sentiment analysis data)

After that, in order to train the model, I had to pick the best classifier to determine which label best fits input given to the model. I went with the SdcaMaximumEntropy classifier, since it's effective for multi-class classification problems like sentiment analysis

Behavior when ignored

This is part of the wider state system, but the behavior when ignored is one of my favorite parts of the project so I think it should get it's own mini-section

If you pick the action "Ignore" three times in a row, the pet AI will recognize that it's being neglected and will perform an action to try to get attention from the user, like spinning around, jumping, or even whimpering.

It decides the action with an Epsilon-Greedy Algorithm, which is a common strategy in RL to balance between exploration (trying new actions) and exploitation (choosing the action with the highest expected reward). The pet AI has a 20% chance to explore (choosing a random action) and an 80% chance to exploit (choosing the action with the highest Q-value)

The Q-values reflect the expected reward for an action. The better the action is considered in the context of the pet's emotional state and its goal of getting the user's attention.

Example demonstration

Interact with your virtual AI pet! Type "Feed", "Play", "Praise", "Ignore", "Yell", "Take away toy", "Scold" or chat with it. Type "exit" to exit the program.
> Feed
Pet reacts to Feed: Reward 3
Pet state: Neutral | Emotion Score: 3
> Play
Pet reacts to Play: Reward 3.18
Pet state: Happy | Emotion Score: 6
> Praise
Pet reacts to Praise: Reward 5.6
Pet state: Happy | Emotion Score: 11
> You're a good doggo :)
Pet detects sentiment: Positive
Positive score: 88.2% | Neutral score: 0.9% | Negative score: 10.9%
Pet state: Happy | Emotion Score: 14

About

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - rosalgie/VirtualPetAI: Pet AI using Q-Learning and Sentiment Analysis in ML.NET · GitHub
Skip to content

Repository files navigation

VirtualPetAI

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Features

  • (Modified) Reinforcement Learning (RL) with Q-Learning: The pet AI learns over time by maximizing future rewards based on user interactions. However, it's not really true Q-Learning because the user still decides the actions. The pet AI learns which actions get the best approval of the user.
  • Sentiment Analysis: The pet AI can interpret the tone of the text you type and changes its state based on it.
  • State System: The pet AI has different states like "Happy", "Sad", "Excited", etc., based on the emotionScore, which is affected by actions taken by the user towards the pet AI

How it works

1. Q-Learning

The Q-learning formula I'm using is:

$Q(s, a) = (1 - α) * Q(s, a) + α * (reward + γ * maxFutureReward)$

Where:

  • $Q(s, a)$ represents the expected future reward for taking action a in state s.
  • α (Alpha) is the learning rate, controlling how much new information influences the Q-value (0 < α ≤ 1).
    • If α is close to 0, the pet AI learns slowly, relying more on past experiences for behavior
    • If α is close to 1, the pet AI learns quickly, relying more on new experiences for behavior
  • reward is the immediate reward for taking action a in state s.
  • γ (Gamma) is the discount factor, determining how much future rewards matter.
    • If γ = 0, only immediate rewards are considered.
    • If γ is close to 1, future rewards are given more importance.
  • maxFutureReward is the best possible reward achievable from the next state.

To be honest, this took the most amount of time for me to wrap my head around since I'm not the best at math, and I still don't fully understand it.

This formula is primarily what decides how the pet AI will react to the user's actions (Feed, Praise, Scold, etc.)

2. Sentiment Analysis

In this project, I'm using ML.NET an open-source machine learning framework.

For my sentiment analysis model, I decided to use text classification, where the model receives an input text (ex: "Bad boy! Stop chewing that!") and is asked to guess the output ("Positive, "Neutral", "Negative").

In order to predict the sentiment of text, I had to somehow find a ton of labeled data to feed the model. I couldn't find much of anything for sentiment analysis data based on humans interacting with pets, so I used synthetic data (basically, I prompted ChatGPT to give me a bunch of labeled sentiment analysis data)

After that, in order to train the model, I had to pick the best classifier to determine which label best fits input given to the model. I went with the SdcaMaximumEntropy classifier, since it's effective for multi-class classification problems like sentiment analysis

Behavior when ignored

This is part of the wider state system, but the behavior when ignored is one of my favorite parts of the project so I think it should get it's own mini-section

If you pick the action "Ignore" three times in a row, the pet AI will recognize that it's being neglected and will perform an action to try to get attention from the user, like spinning around, jumping, or even whimpering.

It decides the action with an Epsilon-Greedy Algorithm, which is a common strategy in RL to balance between exploration (trying new actions) and exploitation (choosing the action with the highest expected reward). The pet AI has a 20% chance to explore (choosing a random action) and an 80% chance to exploit (choosing the action with the highest Q-value)

The Q-values reflect the expected reward for an action. The better the action is considered in the context of the pet's emotional state and its goal of getting the user's attention.

Example demonstration

Interact with your virtual AI pet! Type "Feed", "Play", "Praise", "Ignore", "Yell", "Take away toy", "Scold" or chat with it. Type "exit" to exit the program.
> Feed
Pet reacts to Feed: Reward 3
Pet state: Neutral | Emotion Score: 3
> Play
Pet reacts to Play: Reward 3.18
Pet state: Happy | Emotion Score: 6
> Praise
Pet reacts to Praise: Reward 5.6
Pet state: Happy | Emotion Score: 11
> You're a good doggo :)
Pet detects sentiment: Positive
Positive score: 88.2% | Neutral score: 0.9% | Negative score: 10.9%
Pet state: Happy | Emotion Score: 14

About

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - rosalgie/VirtualPetAI: Pet AI using Q-Learning and Sentiment Analysis in ML.NET · GitHub
Skip to content

Repository files navigation

VirtualPetAI

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Features

  • (Modified) Reinforcement Learning (RL) with Q-Learning: The pet AI learns over time by maximizing future rewards based on user interactions. However, it's not really true Q-Learning because the user still decides the actions. The pet AI learns which actions get the best approval of the user.
  • Sentiment Analysis: The pet AI can interpret the tone of the text you type and changes its state based on it.
  • State System: The pet AI has different states like "Happy", "Sad", "Excited", etc., based on the emotionScore, which is affected by actions taken by the user towards the pet AI

How it works

1. Q-Learning

The Q-learning formula I'm using is:

$Q(s, a) = (1 - α) * Q(s, a) + α * (reward + γ * maxFutureReward)$

Where:

  • $Q(s, a)$ represents the expected future reward for taking action a in state s.
  • α (Alpha) is the learning rate, controlling how much new information influences the Q-value (0 < α ≤ 1).
    • If α is close to 0, the pet AI learns slowly, relying more on past experiences for behavior
    • If α is close to 1, the pet AI learns quickly, relying more on new experiences for behavior
  • reward is the immediate reward for taking action a in state s.
  • γ (Gamma) is the discount factor, determining how much future rewards matter.
    • If γ = 0, only immediate rewards are considered.
    • If γ is close to 1, future rewards are given more importance.
  • maxFutureReward is the best possible reward achievable from the next state.

To be honest, this took the most amount of time for me to wrap my head around since I'm not the best at math, and I still don't fully understand it.

This formula is primarily what decides how the pet AI will react to the user's actions (Feed, Praise, Scold, etc.)

2. Sentiment Analysis

In this project, I'm using ML.NET an open-source machine learning framework.

For my sentiment analysis model, I decided to use text classification, where the model receives an input text (ex: "Bad boy! Stop chewing that!") and is asked to guess the output ("Positive, "Neutral", "Negative").

In order to predict the sentiment of text, I had to somehow find a ton of labeled data to feed the model. I couldn't find much of anything for sentiment analysis data based on humans interacting with pets, so I used synthetic data (basically, I prompted ChatGPT to give me a bunch of labeled sentiment analysis data)

After that, in order to train the model, I had to pick the best classifier to determine which label best fits input given to the model. I went with the SdcaMaximumEntropy classifier, since it's effective for multi-class classification problems like sentiment analysis

Behavior when ignored

This is part of the wider state system, but the behavior when ignored is one of my favorite parts of the project so I think it should get it's own mini-section

If you pick the action "Ignore" three times in a row, the pet AI will recognize that it's being neglected and will perform an action to try to get attention from the user, like spinning around, jumping, or even whimpering.

It decides the action with an Epsilon-Greedy Algorithm, which is a common strategy in RL to balance between exploration (trying new actions) and exploitation (choosing the action with the highest expected reward). The pet AI has a 20% chance to explore (choosing a random action) and an 80% chance to exploit (choosing the action with the highest Q-value)

The Q-values reflect the expected reward for an action. The better the action is considered in the context of the pet's emotional state and its goal of getting the user's attention.

Example demonstration

Interact with your virtual AI pet! Type "Feed", "Play", "Praise", "Ignore", "Yell", "Take away toy", "Scold" or chat with it. Type "exit" to exit the program.
> Feed
Pet reacts to Feed: Reward 3
Pet state: Neutral | Emotion Score: 3
> Play
Pet reacts to Play: Reward 3.18
Pet state: Happy | Emotion Score: 6
> Praise
Pet reacts to Praise: Reward 5.6
Pet state: Happy | Emotion Score: 11
> You're a good doggo :)
Pet detects sentiment: Positive
Positive score: 88.2% | Neutral score: 0.9% | Negative score: 10.9%
Pet state: Happy | Emotion Score: 14

About

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - rosalgie/VirtualPetAI: Pet AI using Q-Learning and Sentiment Analysis in ML.NET · GitHub
Skip to content

Repository files navigation

VirtualPetAI

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Features

  • (Modified) Reinforcement Learning (RL) with Q-Learning: The pet AI learns over time by maximizing future rewards based on user interactions. However, it's not really true Q-Learning because the user still decides the actions. The pet AI learns which actions get the best approval of the user.
  • Sentiment Analysis: The pet AI can interpret the tone of the text you type and changes its state based on it.
  • State System: The pet AI has different states like "Happy", "Sad", "Excited", etc., based on the emotionScore, which is affected by actions taken by the user towards the pet AI

How it works

1. Q-Learning

The Q-learning formula I'm using is:

$Q(s, a) = (1 - α) * Q(s, a) + α * (reward + γ * maxFutureReward)$

Where:

  • $Q(s, a)$ represents the expected future reward for taking action a in state s.
  • α (Alpha) is the learning rate, controlling how much new information influences the Q-value (0 < α ≤ 1).
    • If α is close to 0, the pet AI learns slowly, relying more on past experiences for behavior
    • If α is close to 1, the pet AI learns quickly, relying more on new experiences for behavior
  • reward is the immediate reward for taking action a in state s.
  • γ (Gamma) is the discount factor, determining how much future rewards matter.
    • If γ = 0, only immediate rewards are considered.
    • If γ is close to 1, future rewards are given more importance.
  • maxFutureReward is the best possible reward achievable from the next state.

To be honest, this took the most amount of time for me to wrap my head around since I'm not the best at math, and I still don't fully understand it.

This formula is primarily what decides how the pet AI will react to the user's actions (Feed, Praise, Scold, etc.)

2. Sentiment Analysis

In this project, I'm using ML.NET an open-source machine learning framework.

For my sentiment analysis model, I decided to use text classification, where the model receives an input text (ex: "Bad boy! Stop chewing that!") and is asked to guess the output ("Positive, "Neutral", "Negative").

In order to predict the sentiment of text, I had to somehow find a ton of labeled data to feed the model. I couldn't find much of anything for sentiment analysis data based on humans interacting with pets, so I used synthetic data (basically, I prompted ChatGPT to give me a bunch of labeled sentiment analysis data)

After that, in order to train the model, I had to pick the best classifier to determine which label best fits input given to the model. I went with the SdcaMaximumEntropy classifier, since it's effective for multi-class classification problems like sentiment analysis

Behavior when ignored

This is part of the wider state system, but the behavior when ignored is one of my favorite parts of the project so I think it should get it's own mini-section

If you pick the action "Ignore" three times in a row, the pet AI will recognize that it's being neglected and will perform an action to try to get attention from the user, like spinning around, jumping, or even whimpering.

It decides the action with an Epsilon-Greedy Algorithm, which is a common strategy in RL to balance between exploration (trying new actions) and exploitation (choosing the action with the highest expected reward). The pet AI has a 20% chance to explore (choosing a random action) and an 80% chance to exploit (choosing the action with the highest Q-value)

The Q-values reflect the expected reward for an action. The better the action is considered in the context of the pet's emotional state and its goal of getting the user's attention.

Example demonstration

Interact with your virtual AI pet! Type "Feed", "Play", "Praise", "Ignore", "Yell", "Take away toy", "Scold" or chat with it. Type "exit" to exit the program.
> Feed
Pet reacts to Feed: Reward 3
Pet state: Neutral | Emotion Score: 3
> Play
Pet reacts to Play: Reward 3.18
Pet state: Happy | Emotion Score: 6
> Praise
Pet reacts to Praise: Reward 5.6
Pet state: Happy | Emotion Score: 11
> You're a good doggo :)
Pet detects sentiment: Positive
Positive score: 88.2% | Neutral score: 0.9% | Negative score: 10.9%
Pet state: Happy | Emotion Score: 14

About

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - rosalgie/VirtualPetAI: Pet AI using Q-Learning and Sentiment Analysis in ML.NET · GitHub
Skip to content

Repository files navigation

VirtualPetAI

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Features

  • (Modified) Reinforcement Learning (RL) with Q-Learning: The pet AI learns over time by maximizing future rewards based on user interactions. However, it's not really true Q-Learning because the user still decides the actions. The pet AI learns which actions get the best approval of the user.
  • Sentiment Analysis: The pet AI can interpret the tone of the text you type and changes its state based on it.
  • State System: The pet AI has different states like "Happy", "Sad", "Excited", etc., based on the emotionScore, which is affected by actions taken by the user towards the pet AI

How it works

1. Q-Learning

The Q-learning formula I'm using is:

$Q(s, a) = (1 - α) * Q(s, a) + α * (reward + γ * maxFutureReward)$

Where:

  • $Q(s, a)$ represents the expected future reward for taking action a in state s.
  • α (Alpha) is the learning rate, controlling how much new information influences the Q-value (0 < α ≤ 1).
    • If α is close to 0, the pet AI learns slowly, relying more on past experiences for behavior
    • If α is close to 1, the pet AI learns quickly, relying more on new experiences for behavior
  • reward is the immediate reward for taking action a in state s.
  • γ (Gamma) is the discount factor, determining how much future rewards matter.
    • If γ = 0, only immediate rewards are considered.
    • If γ is close to 1, future rewards are given more importance.
  • maxFutureReward is the best possible reward achievable from the next state.

To be honest, this took the most amount of time for me to wrap my head around since I'm not the best at math, and I still don't fully understand it.

This formula is primarily what decides how the pet AI will react to the user's actions (Feed, Praise, Scold, etc.)

2. Sentiment Analysis

In this project, I'm using ML.NET an open-source machine learning framework.

For my sentiment analysis model, I decided to use text classification, where the model receives an input text (ex: "Bad boy! Stop chewing that!") and is asked to guess the output ("Positive, "Neutral", "Negative").

In order to predict the sentiment of text, I had to somehow find a ton of labeled data to feed the model. I couldn't find much of anything for sentiment analysis data based on humans interacting with pets, so I used synthetic data (basically, I prompted ChatGPT to give me a bunch of labeled sentiment analysis data)

After that, in order to train the model, I had to pick the best classifier to determine which label best fits input given to the model. I went with the SdcaMaximumEntropy classifier, since it's effective for multi-class classification problems like sentiment analysis

Behavior when ignored

This is part of the wider state system, but the behavior when ignored is one of my favorite parts of the project so I think it should get it's own mini-section

If you pick the action "Ignore" three times in a row, the pet AI will recognize that it's being neglected and will perform an action to try to get attention from the user, like spinning around, jumping, or even whimpering.

It decides the action with an Epsilon-Greedy Algorithm, which is a common strategy in RL to balance between exploration (trying new actions) and exploitation (choosing the action with the highest expected reward). The pet AI has a 20% chance to explore (choosing a random action) and an 80% chance to exploit (choosing the action with the highest Q-value)

The Q-values reflect the expected reward for an action. The better the action is considered in the context of the pet's emotional state and its goal of getting the user's attention.

Example demonstration

Interact with your virtual AI pet! Type "Feed", "Play", "Praise", "Ignore", "Yell", "Take away toy", "Scold" or chat with it. Type "exit" to exit the program.
> Feed
Pet reacts to Feed: Reward 3
Pet state: Neutral | Emotion Score: 3
> Play
Pet reacts to Play: Reward 3.18
Pet state: Happy | Emotion Score: 6
> Praise
Pet reacts to Praise: Reward 5.6
Pet state: Happy | Emotion Score: 11
> You're a good doggo :)
Pet detects sentiment: Positive
Positive score: 88.2% | Neutral score: 0.9% | Negative score: 10.9%
Pet state: Happy | Emotion Score: 14

About

Pet AI using Q-Learning and Sentiment Analysis in ML.NET

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages