Skip to content

Repository files navigation

Gem VersionBuild StatusCode ClimateCoverage Status

RubySpeech

RubySpeech is a library for constructing and parsing Text to Speech (TTS) and Automatic Speech Recognition (ASR) documents such as SSML, GRXML and NLSML. Such documents can be constructed to be processed by TTS and ASR engines, parsed as the result from such, or used in the implementation of such engines.

Dependencies

pcre (except on JRuby)

On OSX with Homebrew

brew install pcre

On Ubuntu/Debian

sudo apt-get install libpcre3 libpcre3-dev

On CentOS

sudo yum install pcre-devel

Installation

gem install ruby_speech

Ruby Version Compatability

  • CRuby 2.1+
  • JRuby 9.1+

Library

SSML

RubySpeech provides a DSL for constructing SSML documents like so:

require'ruby_speech'speak=RubySpeech::SSML.drawdovoicegender: :male,name: 'fred'dostring"Hi, I'm Fred. The time is currently "say_asinterpret_as: 'date',format: 'dmy'do"01/02/1960"endendendspeak.to_s

becomes:

<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

Once your Speak is fully prepared and you're ready to send it off for processing, you must call to_doc on it to add the XML header:

<?xml version="1.0"?>
<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

You may also then need to call to_s.

GRXML

Construct a GRXML (SRGS) document like this:

require'ruby_speech'grammy=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendgrammy.to_s

which becomes

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="pin">
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
<ruleid="pin"scope="public">
<one-of>
<item><itemrepeat="4"><rulerefuri="#digit"/></item>#</item>
<item>* 9</item>
</one-of>
</rule>
</grammar>

Built-in grammars

There are some grammars pre-defined which are available from the RubySpeech::GRXML::Builtins module like so:

require'ruby_speech'RubySpeech::GRXML::Builtins.currency

which yields

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="currency">
<ruleid="currency"scope="public">
<itemrepeat="0-">
<rulerefuri="#digit"/>
</item>
<item>*</item>
<itemrepeat="2">
<rulerefuri="#digit"/>
</item>
</rule>
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
</grammar>

These grammars come from the VoiceXML specification, and can be used as indicated there (including parameterisation). They can be used just like any you would manually create, and there's nothing special about them except that they are already defined for you. A full list of available grammars can be found in the API documentation.

These grammars are also available via URI like so:

require'ruby_speech'RubySpeech::GRXML.from_uri('builtin:dtmf/boolean?y=3;n=4')

Grammar matching

It is possible to match some arbitrary input against a GRXML grammar, like so:

require'ruby_speech'
>> grammar=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendmatcher=RubySpeech::GRXML::Matcher.newgrammar
>> matcher.match'*9'=>#<RubySpeech::GRXML::Match:0x00000100ae5d98@mode=:dtmf,@confidence=1,@utterance="*9",@interpretation="*9"
>
>> matcher.match'1234#'=>#<RubySpeech::GRXML::Match:0x00000100b7e020@mode=:dtmf,@confidence=1,@utterance="1234#",@interpretation="1234#"
>
>> matcher.match'5678#'=>#<RubySpeech::GRXML::Match:0x00000101218688@mode=:dtmf,@confidence=1,@utterance="5678#",@interpretation="5678#"
>
>> matcher.match'1111#'=>#<RubySpeech::GRXML::Match:0x000001012f69d8@mode=:dtmf,@confidence=1,@utterance="1111#",@interpretation="1111#"
>
>> matcher.match'111'=>#<RubySpeech::GRXML::NoMatch:0x00000101371660>

NLSML

Natural Language Semantics Markup Language is the format used by many Speech Recognition engines and natural language processors to add semantic information to human language. RubySpeech is capable of generating and parsing such documents.

It is possible to generate an NLSML document like so:

require'ruby_speech'nlsml=RubySpeech::NLSML.drawgrammar: 'http://flight'dointerpretationconfidence: 0.6doinput"I want to go to Pittsburgh",mode: :voiceinstancedoairlinedoto_city'Pittsburgh'endendendinterpretationconfidence: 0.4doinput"I want to go to Stockholm"instancedoairlinedoto_city"Stockholm"endendendendnlsml.to_s

becomes:

<?xml version="1.0"?>
<resultxmlns="http://www.ietf.org/xml/ns/mrcpv2"grammar="http://flight">
<interpretationconfidence="0.6">
<inputmode="voice">I want to go to Pittsburgh</input>
<instance>
<airline>
<to_city>Pittsburgh</to_city>
</airline>
</instance>
</interpretation>
<interpretationconfidence="0.4">
<input>I want to go to Stockholm</input>
<instance>
<airline>
<to_city>Stockholm</to_city>
</airline>
</instance>
</interpretation>
</result>

It's also possible to parse an NLSML document and extract useful information from it. Taking the above example, one may do:

document=RubySpeech.parsenlsml.to_sdocument.match?# => truedocument.interpretations# => [{confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}},{confidence: 0.4,input: {content: 'I want to go to Stockholm'},instance: {airline: {to_city: 'Stockholm'}}}]document.best_interpretation# => {confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}}

Check out the YARD documentation for more

Features:

SSML

  • Document construction
  • <voice/>
  • <prosody/>
  • <emphasis/>
  • <say-as/>
  • <break/>
  • <audio/>
  • <p/> and <s/>
  • <phoneme/>
  • <sub/>

Misc

  • <mark/>
  • <desc/>

GRXML

  • Document construction
  • <item/>
  • <one-of/>
  • <rule/>
  • <ruleref/>
  • <tag/>
  • <token/>

NLSML

  • Document construction
  • Simple data extraction from documents

TODO:

SSML

  • <lexicon/>
  • <meta/> and <metadata/>

GRXML

  • <meta/> and <metadata/>
  • <example/>
  • <lexicon/>

Links:

Note on Patches/Pull Requests

  • Fork the project.
  • Make your feature addition or bug fix.
  • Add tests for it. This is important so I don't break it in a future version unintentionally.
  • Commit, do not mess with rakefile, version, or history.
    • If you want to have your own version, that is fine but bump version in a commit by itself so I can ignore when I pull
  • Send me a pull request. Bonus points for topic branches.

Copyright

Copyright (c) 2013 Ben Langfeld. MIT licence (see LICENSE for details).

About

A ruby library for TTS & ASR document preparation

Resources

Stars

101 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - adhearsion/ruby_speech: A ruby library for TTS & ASR document preparation · GitHub
Skip to content

Repository files navigation

Gem VersionBuild StatusCode ClimateCoverage Status

RubySpeech

RubySpeech is a library for constructing and parsing Text to Speech (TTS) and Automatic Speech Recognition (ASR) documents such as SSML, GRXML and NLSML. Such documents can be constructed to be processed by TTS and ASR engines, parsed as the result from such, or used in the implementation of such engines.

Dependencies

pcre (except on JRuby)

On OSX with Homebrew

brew install pcre

On Ubuntu/Debian

sudo apt-get install libpcre3 libpcre3-dev

On CentOS

sudo yum install pcre-devel

Installation

gem install ruby_speech

Ruby Version Compatability

  • CRuby 2.1+
  • JRuby 9.1+

Library

SSML

RubySpeech provides a DSL for constructing SSML documents like so:

require'ruby_speech'speak=RubySpeech::SSML.drawdovoicegender: :male,name: 'fred'dostring"Hi, I'm Fred. The time is currently "say_asinterpret_as: 'date',format: 'dmy'do"01/02/1960"endendendspeak.to_s

becomes:

<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

Once your Speak is fully prepared and you're ready to send it off for processing, you must call to_doc on it to add the XML header:

<?xml version="1.0"?>
<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

You may also then need to call to_s.

GRXML

Construct a GRXML (SRGS) document like this:

require'ruby_speech'grammy=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendgrammy.to_s

which becomes

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="pin">
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
<ruleid="pin"scope="public">
<one-of>
<item><itemrepeat="4"><rulerefuri="#digit"/></item>#</item>
<item>* 9</item>
</one-of>
</rule>
</grammar>

Built-in grammars

There are some grammars pre-defined which are available from the RubySpeech::GRXML::Builtins module like so:

require'ruby_speech'RubySpeech::GRXML::Builtins.currency

which yields

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="currency">
<ruleid="currency"scope="public">
<itemrepeat="0-">
<rulerefuri="#digit"/>
</item>
<item>*</item>
<itemrepeat="2">
<rulerefuri="#digit"/>
</item>
</rule>
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
</grammar>

These grammars come from the VoiceXML specification, and can be used as indicated there (including parameterisation). They can be used just like any you would manually create, and there's nothing special about them except that they are already defined for you. A full list of available grammars can be found in the API documentation.

These grammars are also available via URI like so:

require'ruby_speech'RubySpeech::GRXML.from_uri('builtin:dtmf/boolean?y=3;n=4')

Grammar matching

It is possible to match some arbitrary input against a GRXML grammar, like so:

require'ruby_speech'
>> grammar=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendmatcher=RubySpeech::GRXML::Matcher.newgrammar
>> matcher.match'*9'=>#<RubySpeech::GRXML::Match:0x00000100ae5d98@mode=:dtmf,@confidence=1,@utterance="*9",@interpretation="*9"
>
>> matcher.match'1234#'=>#<RubySpeech::GRXML::Match:0x00000100b7e020@mode=:dtmf,@confidence=1,@utterance="1234#",@interpretation="1234#"
>
>> matcher.match'5678#'=>#<RubySpeech::GRXML::Match:0x00000101218688@mode=:dtmf,@confidence=1,@utterance="5678#",@interpretation="5678#"
>
>> matcher.match'1111#'=>#<RubySpeech::GRXML::Match:0x000001012f69d8@mode=:dtmf,@confidence=1,@utterance="1111#",@interpretation="1111#"
>
>> matcher.match'111'=>#<RubySpeech::GRXML::NoMatch:0x00000101371660>

NLSML

Natural Language Semantics Markup Language is the format used by many Speech Recognition engines and natural language processors to add semantic information to human language. RubySpeech is capable of generating and parsing such documents.

It is possible to generate an NLSML document like so:

require'ruby_speech'nlsml=RubySpeech::NLSML.drawgrammar: 'http://flight'dointerpretationconfidence: 0.6doinput"I want to go to Pittsburgh",mode: :voiceinstancedoairlinedoto_city'Pittsburgh'endendendinterpretationconfidence: 0.4doinput"I want to go to Stockholm"instancedoairlinedoto_city"Stockholm"endendendendnlsml.to_s

becomes:

<?xml version="1.0"?>
<resultxmlns="http://www.ietf.org/xml/ns/mrcpv2"grammar="http://flight">
<interpretationconfidence="0.6">
<inputmode="voice">I want to go to Pittsburgh</input>
<instance>
<airline>
<to_city>Pittsburgh</to_city>
</airline>
</instance>
</interpretation>
<interpretationconfidence="0.4">
<input>I want to go to Stockholm</input>
<instance>
<airline>
<to_city>Stockholm</to_city>
</airline>
</instance>
</interpretation>
</result>

It's also possible to parse an NLSML document and extract useful information from it. Taking the above example, one may do:

document=RubySpeech.parsenlsml.to_sdocument.match?# => truedocument.interpretations# => [{confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}},{confidence: 0.4,input: {content: 'I want to go to Stockholm'},instance: {airline: {to_city: 'Stockholm'}}}]document.best_interpretation# => {confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}}

Check out the YARD documentation for more

Features:

SSML

  • Document construction
  • <voice/>
  • <prosody/>
  • <emphasis/>
  • <say-as/>
  • <break/>
  • <audio/>
  • <p/> and <s/>
  • <phoneme/>
  • <sub/>

Misc

  • <mark/>
  • <desc/>

GRXML

  • Document construction
  • <item/>
  • <one-of/>
  • <rule/>
  • <ruleref/>
  • <tag/>
  • <token/>

NLSML

  • Document construction
  • Simple data extraction from documents

TODO:

SSML

  • <lexicon/>
  • <meta/> and <metadata/>

GRXML

  • <meta/> and <metadata/>
  • <example/>
  • <lexicon/>

Links:

Note on Patches/Pull Requests

  • Fork the project.
  • Make your feature addition or bug fix.
  • Add tests for it. This is important so I don't break it in a future version unintentionally.
  • Commit, do not mess with rakefile, version, or history.
    • If you want to have your own version, that is fine but bump version in a commit by itself so I can ignore when I pull
  • Send me a pull request. Bonus points for topic branches.

Copyright

Copyright (c) 2013 Ben Langfeld. MIT licence (see LICENSE for details).

About

A ruby library for TTS & ASR document preparation

Resources

Stars

101 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - adhearsion/ruby_speech: A ruby library for TTS & ASR document preparation · GitHub
Skip to content

Repository files navigation

Gem VersionBuild StatusCode ClimateCoverage Status

RubySpeech

RubySpeech is a library for constructing and parsing Text to Speech (TTS) and Automatic Speech Recognition (ASR) documents such as SSML, GRXML and NLSML. Such documents can be constructed to be processed by TTS and ASR engines, parsed as the result from such, or used in the implementation of such engines.

Dependencies

pcre (except on JRuby)

On OSX with Homebrew

brew install pcre

On Ubuntu/Debian

sudo apt-get install libpcre3 libpcre3-dev

On CentOS

sudo yum install pcre-devel

Installation

gem install ruby_speech

Ruby Version Compatability

  • CRuby 2.1+
  • JRuby 9.1+

Library

SSML

RubySpeech provides a DSL for constructing SSML documents like so:

require'ruby_speech'speak=RubySpeech::SSML.drawdovoicegender: :male,name: 'fred'dostring"Hi, I'm Fred. The time is currently "say_asinterpret_as: 'date',format: 'dmy'do"01/02/1960"endendendspeak.to_s

becomes:

<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

Once your Speak is fully prepared and you're ready to send it off for processing, you must call to_doc on it to add the XML header:

<?xml version="1.0"?>
<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

You may also then need to call to_s.

GRXML

Construct a GRXML (SRGS) document like this:

require'ruby_speech'grammy=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendgrammy.to_s

which becomes

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="pin">
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
<ruleid="pin"scope="public">
<one-of>
<item><itemrepeat="4"><rulerefuri="#digit"/></item>#</item>
<item>* 9</item>
</one-of>
</rule>
</grammar>

Built-in grammars

There are some grammars pre-defined which are available from the RubySpeech::GRXML::Builtins module like so:

require'ruby_speech'RubySpeech::GRXML::Builtins.currency

which yields

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="currency">
<ruleid="currency"scope="public">
<itemrepeat="0-">
<rulerefuri="#digit"/>
</item>
<item>*</item>
<itemrepeat="2">
<rulerefuri="#digit"/>
</item>
</rule>
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
</grammar>

These grammars come from the VoiceXML specification, and can be used as indicated there (including parameterisation). They can be used just like any you would manually create, and there's nothing special about them except that they are already defined for you. A full list of available grammars can be found in the API documentation.

These grammars are also available via URI like so:

require'ruby_speech'RubySpeech::GRXML.from_uri('builtin:dtmf/boolean?y=3;n=4')

Grammar matching

It is possible to match some arbitrary input against a GRXML grammar, like so:

require'ruby_speech'
>> grammar=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendmatcher=RubySpeech::GRXML::Matcher.newgrammar
>> matcher.match'*9'=>#<RubySpeech::GRXML::Match:0x00000100ae5d98@mode=:dtmf,@confidence=1,@utterance="*9",@interpretation="*9"
>
>> matcher.match'1234#'=>#<RubySpeech::GRXML::Match:0x00000100b7e020@mode=:dtmf,@confidence=1,@utterance="1234#",@interpretation="1234#"
>
>> matcher.match'5678#'=>#<RubySpeech::GRXML::Match:0x00000101218688@mode=:dtmf,@confidence=1,@utterance="5678#",@interpretation="5678#"
>
>> matcher.match'1111#'=>#<RubySpeech::GRXML::Match:0x000001012f69d8@mode=:dtmf,@confidence=1,@utterance="1111#",@interpretation="1111#"
>
>> matcher.match'111'=>#<RubySpeech::GRXML::NoMatch:0x00000101371660>

NLSML

Natural Language Semantics Markup Language is the format used by many Speech Recognition engines and natural language processors to add semantic information to human language. RubySpeech is capable of generating and parsing such documents.

It is possible to generate an NLSML document like so:

require'ruby_speech'nlsml=RubySpeech::NLSML.drawgrammar: 'http://flight'dointerpretationconfidence: 0.6doinput"I want to go to Pittsburgh",mode: :voiceinstancedoairlinedoto_city'Pittsburgh'endendendinterpretationconfidence: 0.4doinput"I want to go to Stockholm"instancedoairlinedoto_city"Stockholm"endendendendnlsml.to_s

becomes:

<?xml version="1.0"?>
<resultxmlns="http://www.ietf.org/xml/ns/mrcpv2"grammar="http://flight">
<interpretationconfidence="0.6">
<inputmode="voice">I want to go to Pittsburgh</input>
<instance>
<airline>
<to_city>Pittsburgh</to_city>
</airline>
</instance>
</interpretation>
<interpretationconfidence="0.4">
<input>I want to go to Stockholm</input>
<instance>
<airline>
<to_city>Stockholm</to_city>
</airline>
</instance>
</interpretation>
</result>

It's also possible to parse an NLSML document and extract useful information from it. Taking the above example, one may do:

document=RubySpeech.parsenlsml.to_sdocument.match?# => truedocument.interpretations# => [{confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}},{confidence: 0.4,input: {content: 'I want to go to Stockholm'},instance: {airline: {to_city: 'Stockholm'}}}]document.best_interpretation# => {confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}}

Check out the YARD documentation for more

Features:

SSML

  • Document construction
  • <voice/>
  • <prosody/>
  • <emphasis/>
  • <say-as/>
  • <break/>
  • <audio/>
  • <p/> and <s/>
  • <phoneme/>
  • <sub/>

Misc

  • <mark/>
  • <desc/>

GRXML

  • Document construction
  • <item/>
  • <one-of/>
  • <rule/>
  • <ruleref/>
  • <tag/>
  • <token/>

NLSML

  • Document construction
  • Simple data extraction from documents

TODO:

SSML

  • <lexicon/>
  • <meta/> and <metadata/>

GRXML

  • <meta/> and <metadata/>
  • <example/>
  • <lexicon/>

Links:

Note on Patches/Pull Requests

  • Fork the project.
  • Make your feature addition or bug fix.
  • Add tests for it. This is important so I don't break it in a future version unintentionally.
  • Commit, do not mess with rakefile, version, or history.
    • If you want to have your own version, that is fine but bump version in a commit by itself so I can ignore when I pull
  • Send me a pull request. Bonus points for topic branches.

Copyright

Copyright (c) 2013 Ben Langfeld. MIT licence (see LICENSE for details).

About

A ruby library for TTS & ASR document preparation

Resources

Stars

101 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - adhearsion/ruby_speech: A ruby library for TTS & ASR document preparation · GitHub
Skip to content

Repository files navigation

Gem VersionBuild StatusCode ClimateCoverage Status

RubySpeech

RubySpeech is a library for constructing and parsing Text to Speech (TTS) and Automatic Speech Recognition (ASR) documents such as SSML, GRXML and NLSML. Such documents can be constructed to be processed by TTS and ASR engines, parsed as the result from such, or used in the implementation of such engines.

Dependencies

pcre (except on JRuby)

On OSX with Homebrew

brew install pcre

On Ubuntu/Debian

sudo apt-get install libpcre3 libpcre3-dev

On CentOS

sudo yum install pcre-devel

Installation

gem install ruby_speech

Ruby Version Compatability

  • CRuby 2.1+
  • JRuby 9.1+

Library

SSML

RubySpeech provides a DSL for constructing SSML documents like so:

require'ruby_speech'speak=RubySpeech::SSML.drawdovoicegender: :male,name: 'fred'dostring"Hi, I'm Fred. The time is currently "say_asinterpret_as: 'date',format: 'dmy'do"01/02/1960"endendendspeak.to_s

becomes:

<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

Once your Speak is fully prepared and you're ready to send it off for processing, you must call to_doc on it to add the XML header:

<?xml version="1.0"?>
<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

You may also then need to call to_s.

GRXML

Construct a GRXML (SRGS) document like this:

require'ruby_speech'grammy=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendgrammy.to_s

which becomes

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="pin">
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
<ruleid="pin"scope="public">
<one-of>
<item><itemrepeat="4"><rulerefuri="#digit"/></item>#</item>
<item>* 9</item>
</one-of>
</rule>
</grammar>

Built-in grammars

There are some grammars pre-defined which are available from the RubySpeech::GRXML::Builtins module like so:

require'ruby_speech'RubySpeech::GRXML::Builtins.currency

which yields

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="currency">
<ruleid="currency"scope="public">
<itemrepeat="0-">
<rulerefuri="#digit"/>
</item>
<item>*</item>
<itemrepeat="2">
<rulerefuri="#digit"/>
</item>
</rule>
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
</grammar>

These grammars come from the VoiceXML specification, and can be used as indicated there (including parameterisation). They can be used just like any you would manually create, and there's nothing special about them except that they are already defined for you. A full list of available grammars can be found in the API documentation.

These grammars are also available via URI like so:

require'ruby_speech'RubySpeech::GRXML.from_uri('builtin:dtmf/boolean?y=3;n=4')

Grammar matching

It is possible to match some arbitrary input against a GRXML grammar, like so:

require'ruby_speech'
>> grammar=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendmatcher=RubySpeech::GRXML::Matcher.newgrammar
>> matcher.match'*9'=>#<RubySpeech::GRXML::Match:0x00000100ae5d98@mode=:dtmf,@confidence=1,@utterance="*9",@interpretation="*9"
>
>> matcher.match'1234#'=>#<RubySpeech::GRXML::Match:0x00000100b7e020@mode=:dtmf,@confidence=1,@utterance="1234#",@interpretation="1234#"
>
>> matcher.match'5678#'=>#<RubySpeech::GRXML::Match:0x00000101218688@mode=:dtmf,@confidence=1,@utterance="5678#",@interpretation="5678#"
>
>> matcher.match'1111#'=>#<RubySpeech::GRXML::Match:0x000001012f69d8@mode=:dtmf,@confidence=1,@utterance="1111#",@interpretation="1111#"
>
>> matcher.match'111'=>#<RubySpeech::GRXML::NoMatch:0x00000101371660>

NLSML

Natural Language Semantics Markup Language is the format used by many Speech Recognition engines and natural language processors to add semantic information to human language. RubySpeech is capable of generating and parsing such documents.

It is possible to generate an NLSML document like so:

require'ruby_speech'nlsml=RubySpeech::NLSML.drawgrammar: 'http://flight'dointerpretationconfidence: 0.6doinput"I want to go to Pittsburgh",mode: :voiceinstancedoairlinedoto_city'Pittsburgh'endendendinterpretationconfidence: 0.4doinput"I want to go to Stockholm"instancedoairlinedoto_city"Stockholm"endendendendnlsml.to_s

becomes:

<?xml version="1.0"?>
<resultxmlns="http://www.ietf.org/xml/ns/mrcpv2"grammar="http://flight">
<interpretationconfidence="0.6">
<inputmode="voice">I want to go to Pittsburgh</input>
<instance>
<airline>
<to_city>Pittsburgh</to_city>
</airline>
</instance>
</interpretation>
<interpretationconfidence="0.4">
<input>I want to go to Stockholm</input>
<instance>
<airline>
<to_city>Stockholm</to_city>
</airline>
</instance>
</interpretation>
</result>

It's also possible to parse an NLSML document and extract useful information from it. Taking the above example, one may do:

document=RubySpeech.parsenlsml.to_sdocument.match?# => truedocument.interpretations# => [{confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}},{confidence: 0.4,input: {content: 'I want to go to Stockholm'},instance: {airline: {to_city: 'Stockholm'}}}]document.best_interpretation# => {confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}}

Check out the YARD documentation for more

Features:

SSML

  • Document construction
  • <voice/>
  • <prosody/>
  • <emphasis/>
  • <say-as/>
  • <break/>
  • <audio/>
  • <p/> and <s/>
  • <phoneme/>
  • <sub/>

Misc

  • <mark/>
  • <desc/>

GRXML

  • Document construction
  • <item/>
  • <one-of/>
  • <rule/>
  • <ruleref/>
  • <tag/>
  • <token/>

NLSML

  • Document construction
  • Simple data extraction from documents

TODO:

SSML

  • <lexicon/>
  • <meta/> and <metadata/>

GRXML

  • <meta/> and <metadata/>
  • <example/>
  • <lexicon/>

Links:

Note on Patches/Pull Requests

  • Fork the project.
  • Make your feature addition or bug fix.
  • Add tests for it. This is important so I don't break it in a future version unintentionally.
  • Commit, do not mess with rakefile, version, or history.
    • If you want to have your own version, that is fine but bump version in a commit by itself so I can ignore when I pull
  • Send me a pull request. Bonus points for topic branches.

Copyright

Copyright (c) 2013 Ben Langfeld. MIT licence (see LICENSE for details).

About

A ruby library for TTS & ASR document preparation

Resources

Stars

101 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - adhearsion/ruby_speech: A ruby library for TTS & ASR document preparation · GitHub
Skip to content

Repository files navigation

Gem VersionBuild StatusCode ClimateCoverage Status

RubySpeech

RubySpeech is a library for constructing and parsing Text to Speech (TTS) and Automatic Speech Recognition (ASR) documents such as SSML, GRXML and NLSML. Such documents can be constructed to be processed by TTS and ASR engines, parsed as the result from such, or used in the implementation of such engines.

Dependencies

pcre (except on JRuby)

On OSX with Homebrew

brew install pcre

On Ubuntu/Debian

sudo apt-get install libpcre3 libpcre3-dev

On CentOS

sudo yum install pcre-devel

Installation

gem install ruby_speech

Ruby Version Compatability

  • CRuby 2.1+
  • JRuby 9.1+

Library

SSML

RubySpeech provides a DSL for constructing SSML documents like so:

require'ruby_speech'speak=RubySpeech::SSML.drawdovoicegender: :male,name: 'fred'dostring"Hi, I'm Fred. The time is currently "say_asinterpret_as: 'date',format: 'dmy'do"01/02/1960"endendendspeak.to_s

becomes:

<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

Once your Speak is fully prepared and you're ready to send it off for processing, you must call to_doc on it to add the XML header:

<?xml version="1.0"?>
<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

You may also then need to call to_s.

GRXML

Construct a GRXML (SRGS) document like this:

require'ruby_speech'grammy=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendgrammy.to_s

which becomes

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="pin">
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
<ruleid="pin"scope="public">
<one-of>
<item><itemrepeat="4"><rulerefuri="#digit"/></item>#</item>
<item>* 9</item>
</one-of>
</rule>
</grammar>

Built-in grammars

There are some grammars pre-defined which are available from the RubySpeech::GRXML::Builtins module like so:

require'ruby_speech'RubySpeech::GRXML::Builtins.currency

which yields

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="currency">
<ruleid="currency"scope="public">
<itemrepeat="0-">
<rulerefuri="#digit"/>
</item>
<item>*</item>
<itemrepeat="2">
<rulerefuri="#digit"/>
</item>
</rule>
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
</grammar>

These grammars come from the VoiceXML specification, and can be used as indicated there (including parameterisation). They can be used just like any you would manually create, and there's nothing special about them except that they are already defined for you. A full list of available grammars can be found in the API documentation.

These grammars are also available via URI like so:

require'ruby_speech'RubySpeech::GRXML.from_uri('builtin:dtmf/boolean?y=3;n=4')

Grammar matching

It is possible to match some arbitrary input against a GRXML grammar, like so:

require'ruby_speech'
>> grammar=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendmatcher=RubySpeech::GRXML::Matcher.newgrammar
>> matcher.match'*9'=>#<RubySpeech::GRXML::Match:0x00000100ae5d98@mode=:dtmf,@confidence=1,@utterance="*9",@interpretation="*9"
>
>> matcher.match'1234#'=>#<RubySpeech::GRXML::Match:0x00000100b7e020@mode=:dtmf,@confidence=1,@utterance="1234#",@interpretation="1234#"
>
>> matcher.match'5678#'=>#<RubySpeech::GRXML::Match:0x00000101218688@mode=:dtmf,@confidence=1,@utterance="5678#",@interpretation="5678#"
>
>> matcher.match'1111#'=>#<RubySpeech::GRXML::Match:0x000001012f69d8@mode=:dtmf,@confidence=1,@utterance="1111#",@interpretation="1111#"
>
>> matcher.match'111'=>#<RubySpeech::GRXML::NoMatch:0x00000101371660>

NLSML

Natural Language Semantics Markup Language is the format used by many Speech Recognition engines and natural language processors to add semantic information to human language. RubySpeech is capable of generating and parsing such documents.

It is possible to generate an NLSML document like so:

require'ruby_speech'nlsml=RubySpeech::NLSML.drawgrammar: 'http://flight'dointerpretationconfidence: 0.6doinput"I want to go to Pittsburgh",mode: :voiceinstancedoairlinedoto_city'Pittsburgh'endendendinterpretationconfidence: 0.4doinput"I want to go to Stockholm"instancedoairlinedoto_city"Stockholm"endendendendnlsml.to_s

becomes:

<?xml version="1.0"?>
<resultxmlns="http://www.ietf.org/xml/ns/mrcpv2"grammar="http://flight">
<interpretationconfidence="0.6">
<inputmode="voice">I want to go to Pittsburgh</input>
<instance>
<airline>
<to_city>Pittsburgh</to_city>
</airline>
</instance>
</interpretation>
<interpretationconfidence="0.4">
<input>I want to go to Stockholm</input>
<instance>
<airline>
<to_city>Stockholm</to_city>
</airline>
</instance>
</interpretation>
</result>

It's also possible to parse an NLSML document and extract useful information from it. Taking the above example, one may do:

document=RubySpeech.parsenlsml.to_sdocument.match?# => truedocument.interpretations# => [{confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}},{confidence: 0.4,input: {content: 'I want to go to Stockholm'},instance: {airline: {to_city: 'Stockholm'}}}]document.best_interpretation# => {confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}}

Check out the YARD documentation for more

Features:

SSML

  • Document construction
  • <voice/>
  • <prosody/>
  • <emphasis/>
  • <say-as/>
  • <break/>
  • <audio/>
  • <p/> and <s/>
  • <phoneme/>
  • <sub/>

Misc

  • <mark/>
  • <desc/>

GRXML

  • Document construction
  • <item/>
  • <one-of/>
  • <rule/>
  • <ruleref/>
  • <tag/>
  • <token/>

NLSML

  • Document construction
  • Simple data extraction from documents

TODO:

SSML

  • <lexicon/>
  • <meta/> and <metadata/>

GRXML

  • <meta/> and <metadata/>
  • <example/>
  • <lexicon/>

Links:

Note on Patches/Pull Requests

  • Fork the project.
  • Make your feature addition or bug fix.
  • Add tests for it. This is important so I don't break it in a future version unintentionally.
  • Commit, do not mess with rakefile, version, or history.
    • If you want to have your own version, that is fine but bump version in a commit by itself so I can ignore when I pull
  • Send me a pull request. Bonus points for topic branches.

Copyright

Copyright (c) 2013 Ben Langfeld. MIT licence (see LICENSE for details).

About

A ruby library for TTS & ASR document preparation

Resources

Stars

101 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - adhearsion/ruby_speech: A ruby library for TTS & ASR document preparation · GitHub
Skip to content

Repository files navigation

Gem VersionBuild StatusCode ClimateCoverage Status

RubySpeech

RubySpeech is a library for constructing and parsing Text to Speech (TTS) and Automatic Speech Recognition (ASR) documents such as SSML, GRXML and NLSML. Such documents can be constructed to be processed by TTS and ASR engines, parsed as the result from such, or used in the implementation of such engines.

Dependencies

pcre (except on JRuby)

On OSX with Homebrew

brew install pcre

On Ubuntu/Debian

sudo apt-get install libpcre3 libpcre3-dev

On CentOS

sudo yum install pcre-devel

Installation

gem install ruby_speech

Ruby Version Compatability

  • CRuby 2.1+
  • JRuby 9.1+

Library

SSML

RubySpeech provides a DSL for constructing SSML documents like so:

require'ruby_speech'speak=RubySpeech::SSML.drawdovoicegender: :male,name: 'fred'dostring"Hi, I'm Fred. The time is currently "say_asinterpret_as: 'date',format: 'dmy'do"01/02/1960"endendendspeak.to_s

becomes:

<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

Once your Speak is fully prepared and you're ready to send it off for processing, you must call to_doc on it to add the XML header:

<?xml version="1.0"?>
<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

You may also then need to call to_s.

GRXML

Construct a GRXML (SRGS) document like this:

require'ruby_speech'grammy=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendgrammy.to_s

which becomes

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="pin">
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
<ruleid="pin"scope="public">
<one-of>
<item><itemrepeat="4"><rulerefuri="#digit"/></item>#</item>
<item>* 9</item>
</one-of>
</rule>
</grammar>

Built-in grammars

There are some grammars pre-defined which are available from the RubySpeech::GRXML::Builtins module like so:

require'ruby_speech'RubySpeech::GRXML::Builtins.currency

which yields

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="currency">
<ruleid="currency"scope="public">
<itemrepeat="0-">
<rulerefuri="#digit"/>
</item>
<item>*</item>
<itemrepeat="2">
<rulerefuri="#digit"/>
</item>
</rule>
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
</grammar>

These grammars come from the VoiceXML specification, and can be used as indicated there (including parameterisation). They can be used just like any you would manually create, and there's nothing special about them except that they are already defined for you. A full list of available grammars can be found in the API documentation.

These grammars are also available via URI like so:

require'ruby_speech'RubySpeech::GRXML.from_uri('builtin:dtmf/boolean?y=3;n=4')

Grammar matching

It is possible to match some arbitrary input against a GRXML grammar, like so:

require'ruby_speech'
>> grammar=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendmatcher=RubySpeech::GRXML::Matcher.newgrammar
>> matcher.match'*9'=>#<RubySpeech::GRXML::Match:0x00000100ae5d98@mode=:dtmf,@confidence=1,@utterance="*9",@interpretation="*9"
>
>> matcher.match'1234#'=>#<RubySpeech::GRXML::Match:0x00000100b7e020@mode=:dtmf,@confidence=1,@utterance="1234#",@interpretation="1234#"
>
>> matcher.match'5678#'=>#<RubySpeech::GRXML::Match:0x00000101218688@mode=:dtmf,@confidence=1,@utterance="5678#",@interpretation="5678#"
>
>> matcher.match'1111#'=>#<RubySpeech::GRXML::Match:0x000001012f69d8@mode=:dtmf,@confidence=1,@utterance="1111#",@interpretation="1111#"
>
>> matcher.match'111'=>#<RubySpeech::GRXML::NoMatch:0x00000101371660>

NLSML

Natural Language Semantics Markup Language is the format used by many Speech Recognition engines and natural language processors to add semantic information to human language. RubySpeech is capable of generating and parsing such documents.

It is possible to generate an NLSML document like so:

require'ruby_speech'nlsml=RubySpeech::NLSML.drawgrammar: 'http://flight'dointerpretationconfidence: 0.6doinput"I want to go to Pittsburgh",mode: :voiceinstancedoairlinedoto_city'Pittsburgh'endendendinterpretationconfidence: 0.4doinput"I want to go to Stockholm"instancedoairlinedoto_city"Stockholm"endendendendnlsml.to_s

becomes:

<?xml version="1.0"?>
<resultxmlns="http://www.ietf.org/xml/ns/mrcpv2"grammar="http://flight">
<interpretationconfidence="0.6">
<inputmode="voice">I want to go to Pittsburgh</input>
<instance>
<airline>
<to_city>Pittsburgh</to_city>
</airline>
</instance>
</interpretation>
<interpretationconfidence="0.4">
<input>I want to go to Stockholm</input>
<instance>
<airline>
<to_city>Stockholm</to_city>
</airline>
</instance>
</interpretation>
</result>

It's also possible to parse an NLSML document and extract useful information from it. Taking the above example, one may do:

document=RubySpeech.parsenlsml.to_sdocument.match?# => truedocument.interpretations# => [{confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}},{confidence: 0.4,input: {content: 'I want to go to Stockholm'},instance: {airline: {to_city: 'Stockholm'}}}]document.best_interpretation# => {confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}}

Check out the YARD documentation for more

Features:

SSML

  • Document construction
  • <voice/>
  • <prosody/>
  • <emphasis/>
  • <say-as/>
  • <break/>
  • <audio/>
  • <p/> and <s/>
  • <phoneme/>
  • <sub/>

Misc

  • <mark/>
  • <desc/>

GRXML

  • Document construction
  • <item/>
  • <one-of/>
  • <rule/>
  • <ruleref/>
  • <tag/>
  • <token/>

NLSML

  • Document construction
  • Simple data extraction from documents

TODO:

SSML

  • <lexicon/>
  • <meta/> and <metadata/>

GRXML

  • <meta/> and <metadata/>
  • <example/>
  • <lexicon/>

Links:

Note on Patches/Pull Requests

  • Fork the project.
  • Make your feature addition or bug fix.
  • Add tests for it. This is important so I don't break it in a future version unintentionally.
  • Commit, do not mess with rakefile, version, or history.
    • If you want to have your own version, that is fine but bump version in a commit by itself so I can ignore when I pull
  • Send me a pull request. Bonus points for topic branches.

Copyright

Copyright (c) 2013 Ben Langfeld. MIT licence (see LICENSE for details).

About

A ruby library for TTS & ASR document preparation

Resources

Stars

101 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - adhearsion/ruby_speech: A ruby library for TTS & ASR document preparation · GitHub
Skip to content

Repository files navigation

Gem VersionBuild StatusCode ClimateCoverage Status

RubySpeech

RubySpeech is a library for constructing and parsing Text to Speech (TTS) and Automatic Speech Recognition (ASR) documents such as SSML, GRXML and NLSML. Such documents can be constructed to be processed by TTS and ASR engines, parsed as the result from such, or used in the implementation of such engines.

Dependencies

pcre (except on JRuby)

On OSX with Homebrew

brew install pcre

On Ubuntu/Debian

sudo apt-get install libpcre3 libpcre3-dev

On CentOS

sudo yum install pcre-devel

Installation

gem install ruby_speech

Ruby Version Compatability

  • CRuby 2.1+
  • JRuby 9.1+

Library

SSML

RubySpeech provides a DSL for constructing SSML documents like so:

require'ruby_speech'speak=RubySpeech::SSML.drawdovoicegender: :male,name: 'fred'dostring"Hi, I'm Fred. The time is currently "say_asinterpret_as: 'date',format: 'dmy'do"01/02/1960"endendendspeak.to_s

becomes:

<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

Once your Speak is fully prepared and you're ready to send it off for processing, you must call to_doc on it to add the XML header:

<?xml version="1.0"?>
<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

You may also then need to call to_s.

GRXML

Construct a GRXML (SRGS) document like this:

require'ruby_speech'grammy=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendgrammy.to_s

which becomes

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="pin">
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
<ruleid="pin"scope="public">
<one-of>
<item><itemrepeat="4"><rulerefuri="#digit"/></item>#</item>
<item>* 9</item>
</one-of>
</rule>
</grammar>

Built-in grammars

There are some grammars pre-defined which are available from the RubySpeech::GRXML::Builtins module like so:

require'ruby_speech'RubySpeech::GRXML::Builtins.currency

which yields

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="currency">
<ruleid="currency"scope="public">
<itemrepeat="0-">
<rulerefuri="#digit"/>
</item>
<item>*</item>
<itemrepeat="2">
<rulerefuri="#digit"/>
</item>
</rule>
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
</grammar>

These grammars come from the VoiceXML specification, and can be used as indicated there (including parameterisation). They can be used just like any you would manually create, and there's nothing special about them except that they are already defined for you. A full list of available grammars can be found in the API documentation.

These grammars are also available via URI like so:

require'ruby_speech'RubySpeech::GRXML.from_uri('builtin:dtmf/boolean?y=3;n=4')

Grammar matching

It is possible to match some arbitrary input against a GRXML grammar, like so:

require'ruby_speech'
>> grammar=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendmatcher=RubySpeech::GRXML::Matcher.newgrammar
>> matcher.match'*9'=>#<RubySpeech::GRXML::Match:0x00000100ae5d98@mode=:dtmf,@confidence=1,@utterance="*9",@interpretation="*9"
>
>> matcher.match'1234#'=>#<RubySpeech::GRXML::Match:0x00000100b7e020@mode=:dtmf,@confidence=1,@utterance="1234#",@interpretation="1234#"
>
>> matcher.match'5678#'=>#<RubySpeech::GRXML::Match:0x00000101218688@mode=:dtmf,@confidence=1,@utterance="5678#",@interpretation="5678#"
>
>> matcher.match'1111#'=>#<RubySpeech::GRXML::Match:0x000001012f69d8@mode=:dtmf,@confidence=1,@utterance="1111#",@interpretation="1111#"
>
>> matcher.match'111'=>#<RubySpeech::GRXML::NoMatch:0x00000101371660>

NLSML

Natural Language Semantics Markup Language is the format used by many Speech Recognition engines and natural language processors to add semantic information to human language. RubySpeech is capable of generating and parsing such documents.

It is possible to generate an NLSML document like so:

require'ruby_speech'nlsml=RubySpeech::NLSML.drawgrammar: 'http://flight'dointerpretationconfidence: 0.6doinput"I want to go to Pittsburgh",mode: :voiceinstancedoairlinedoto_city'Pittsburgh'endendendinterpretationconfidence: 0.4doinput"I want to go to Stockholm"instancedoairlinedoto_city"Stockholm"endendendendnlsml.to_s

becomes:

<?xml version="1.0"?>
<resultxmlns="http://www.ietf.org/xml/ns/mrcpv2"grammar="http://flight">
<interpretationconfidence="0.6">
<inputmode="voice">I want to go to Pittsburgh</input>
<instance>
<airline>
<to_city>Pittsburgh</to_city>
</airline>
</instance>
</interpretation>
<interpretationconfidence="0.4">
<input>I want to go to Stockholm</input>
<instance>
<airline>
<to_city>Stockholm</to_city>
</airline>
</instance>
</interpretation>
</result>

It's also possible to parse an NLSML document and extract useful information from it. Taking the above example, one may do:

document=RubySpeech.parsenlsml.to_sdocument.match?# => truedocument.interpretations# => [{confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}},{confidence: 0.4,input: {content: 'I want to go to Stockholm'},instance: {airline: {to_city: 'Stockholm'}}}]document.best_interpretation# => {confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}}

Check out the YARD documentation for more

Features:

SSML

  • Document construction
  • <voice/>
  • <prosody/>
  • <emphasis/>
  • <say-as/>
  • <break/>
  • <audio/>
  • <p/> and <s/>
  • <phoneme/>
  • <sub/>

Misc

  • <mark/>
  • <desc/>

GRXML

  • Document construction
  • <item/>
  • <one-of/>
  • <rule/>
  • <ruleref/>
  • <tag/>
  • <token/>

NLSML

  • Document construction
  • Simple data extraction from documents

TODO:

SSML

  • <lexicon/>
  • <meta/> and <metadata/>

GRXML

  • <meta/> and <metadata/>
  • <example/>
  • <lexicon/>

Links:

Note on Patches/Pull Requests

  • Fork the project.
  • Make your feature addition or bug fix.
  • Add tests for it. This is important so I don't break it in a future version unintentionally.
  • Commit, do not mess with rakefile, version, or history.
    • If you want to have your own version, that is fine but bump version in a commit by itself so I can ignore when I pull
  • Send me a pull request. Bonus points for topic branches.

Copyright

Copyright (c) 2013 Ben Langfeld. MIT licence (see LICENSE for details).

About

A ruby library for TTS & ASR document preparation

Resources

Stars

101 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - adhearsion/ruby_speech: A ruby library for TTS & ASR document preparation · GitHub
Skip to content

Repository files navigation

Gem VersionBuild StatusCode ClimateCoverage Status

RubySpeech

RubySpeech is a library for constructing and parsing Text to Speech (TTS) and Automatic Speech Recognition (ASR) documents such as SSML, GRXML and NLSML. Such documents can be constructed to be processed by TTS and ASR engines, parsed as the result from such, or used in the implementation of such engines.

Dependencies

pcre (except on JRuby)

On OSX with Homebrew

brew install pcre

On Ubuntu/Debian

sudo apt-get install libpcre3 libpcre3-dev

On CentOS

sudo yum install pcre-devel

Installation

gem install ruby_speech

Ruby Version Compatability

  • CRuby 2.1+
  • JRuby 9.1+

Library

SSML

RubySpeech provides a DSL for constructing SSML documents like so:

require'ruby_speech'speak=RubySpeech::SSML.drawdovoicegender: :male,name: 'fred'dostring"Hi, I'm Fred. The time is currently "say_asinterpret_as: 'date',format: 'dmy'do"01/02/1960"endendendspeak.to_s

becomes:

<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

Once your Speak is fully prepared and you're ready to send it off for processing, you must call to_doc on it to add the XML header:

<?xml version="1.0"?>
<speakxmlns="http://www.w3.org/2001/10/synthesis"version="1.0"xml:lang="en-US">
<voicegender="male"name="fred">
Hi, I'm Fred. The time is currently <say-asformat="dmy"interpret-as="date">01/02/1960</say-as>
</voice>
</speak>

You may also then need to call to_s.

GRXML

Construct a GRXML (SRGS) document like this:

require'ruby_speech'grammy=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendgrammy.to_s

which becomes

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="pin">
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
<ruleid="pin"scope="public">
<one-of>
<item><itemrepeat="4"><rulerefuri="#digit"/></item>#</item>
<item>* 9</item>
</one-of>
</rule>
</grammar>

Built-in grammars

There are some grammars pre-defined which are available from the RubySpeech::GRXML::Builtins module like so:

require'ruby_speech'RubySpeech::GRXML::Builtins.currency

which yields

<grammarxmlns="http://www.w3.org/2001/06/grammar"version="1.0"xml:lang="en-US"mode="dtmf"root="currency">
<ruleid="currency"scope="public">
<itemrepeat="0-">
<rulerefuri="#digit"/>
</item>
<item>*</item>
<itemrepeat="2">
<rulerefuri="#digit"/>
</item>
</rule>
<ruleid="digit">
<one-of>
<item>0</item>
<item>1</item>
<item>2</item>
<item>3</item>
<item>4</item>
<item>5</item>
<item>6</item>
<item>7</item>
<item>8</item>
<item>9</item>
</one-of>
</rule>
</grammar>

These grammars come from the VoiceXML specification, and can be used as indicated there (including parameterisation). They can be used just like any you would manually create, and there's nothing special about them except that they are already defined for you. A full list of available grammars can be found in the API documentation.

These grammars are also available via URI like so:

require'ruby_speech'RubySpeech::GRXML.from_uri('builtin:dtmf/boolean?y=3;n=4')

Grammar matching

It is possible to match some arbitrary input against a GRXML grammar, like so:

require'ruby_speech'
>> grammar=RubySpeech::GRXML.drawmode: :dtmf,root: 'pin'doruleid: 'digit'doone_ofdo('0'..'9').map{ |d| item{d}}endendruleid: 'pin',scope: 'public'doone_ofdoitemdoitemrepeat: '4'dorulerefuri: '#digit'end"#"enditemdo"* 9"endendendendmatcher=RubySpeech::GRXML::Matcher.newgrammar
>> matcher.match'*9'=>#<RubySpeech::GRXML::Match:0x00000100ae5d98@mode=:dtmf,@confidence=1,@utterance="*9",@interpretation="*9"
>
>> matcher.match'1234#'=>#<RubySpeech::GRXML::Match:0x00000100b7e020@mode=:dtmf,@confidence=1,@utterance="1234#",@interpretation="1234#"
>
>> matcher.match'5678#'=>#<RubySpeech::GRXML::Match:0x00000101218688@mode=:dtmf,@confidence=1,@utterance="5678#",@interpretation="5678#"
>
>> matcher.match'1111#'=>#<RubySpeech::GRXML::Match:0x000001012f69d8@mode=:dtmf,@confidence=1,@utterance="1111#",@interpretation="1111#"
>
>> matcher.match'111'=>#<RubySpeech::GRXML::NoMatch:0x00000101371660>

NLSML

Natural Language Semantics Markup Language is the format used by many Speech Recognition engines and natural language processors to add semantic information to human language. RubySpeech is capable of generating and parsing such documents.

It is possible to generate an NLSML document like so:

require'ruby_speech'nlsml=RubySpeech::NLSML.drawgrammar: 'http://flight'dointerpretationconfidence: 0.6doinput"I want to go to Pittsburgh",mode: :voiceinstancedoairlinedoto_city'Pittsburgh'endendendinterpretationconfidence: 0.4doinput"I want to go to Stockholm"instancedoairlinedoto_city"Stockholm"endendendendnlsml.to_s

becomes:

<?xml version="1.0"?>
<resultxmlns="http://www.ietf.org/xml/ns/mrcpv2"grammar="http://flight">
<interpretationconfidence="0.6">
<inputmode="voice">I want to go to Pittsburgh</input>
<instance>
<airline>
<to_city>Pittsburgh</to_city>
</airline>
</instance>
</interpretation>
<interpretationconfidence="0.4">
<input>I want to go to Stockholm</input>
<instance>
<airline>
<to_city>Stockholm</to_city>
</airline>
</instance>
</interpretation>
</result>

It's also possible to parse an NLSML document and extract useful information from it. Taking the above example, one may do:

document=RubySpeech.parsenlsml.to_sdocument.match?# => truedocument.interpretations# => [{confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}},{confidence: 0.4,input: {content: 'I want to go to Stockholm'},instance: {airline: {to_city: 'Stockholm'}}}]document.best_interpretation# => {confidence: 0.6,input: {mode: :voice,content: 'I want to go to Pittsburgh'},instance: {airline: {to_city: 'Pittsburgh'}}}

Check out the YARD documentation for more

Features:

SSML

  • Document construction
  • <voice/>
  • <prosody/>
  • <emphasis/>
  • <say-as/>
  • <break/>
  • <audio/>
  • <p/> and <s/>
  • <phoneme/>
  • <sub/>

Misc

  • <mark/>
  • <desc/>

GRXML

  • Document construction
  • <item/>
  • <one-of/>
  • <rule/>
  • <ruleref/>
  • <tag/>
  • <token/>

NLSML

  • Document construction
  • Simple data extraction from documents

TODO:

SSML

  • <lexicon/>
  • <meta/> and <metadata/>

GRXML

  • <meta/> and <metadata/>
  • <example/>
  • <lexicon/>

Links:

Note on Patches/Pull Requests

  • Fork the project.
  • Make your feature addition or bug fix.
  • Add tests for it. This is important so I don't break it in a future version unintentionally.
  • Commit, do not mess with rakefile, version, or history.
    • If you want to have your own version, that is fine but bump version in a commit by itself so I can ignore when I pull
  • Send me a pull request. Bonus points for topic branches.

Copyright

Copyright (c) 2013 Ben Langfeld. MIT licence (see LICENSE for details).

About

A ruby library for TTS & ASR document preparation

Resources

Stars

101 stars

Watchers

13 watching

Forks

Releases

Packages

Used by

Contributors

Languages