Repository files navigation

BuildcodecovRuff

PyPIPython VersionLicense

construct-tracker

Track and measure constructs, concepts or categories in text documents. Build interpretable lexicon models quickly by using LLMs. Built on top of the OpenRouterAI package so you can use most Generative AI models.

Why build lexicons?

They can be used to build models that are:

  • interpretable: understand why the model outputs a given score, which can help avoid biases and guarantee the model will detect certain phrases (important for high-risk scenarios to use in tandem with LLMs)
  • lightweight: no GPU needed (unlike LLMs)
  • private and free: you can run on your local computer instead of submitting to a cloud API (OpenAI) which may not be secure
  • have high content validity: measure what you actually want to measure (unlike existing lexicons or models that measure something only slightly related)

If you use, please cite

Low DM, Rankin O, Coppersmith DDL, Bentley KH, Nock MK, Ghosh SS (2024). Using Generative AI to create lexicons for interpretable text models with high content validity. PsyarXiv.


Installation

pip install construct-tracker

Measure 49 suicide risk factors in text data

Highlight matches

TutorialOpen in Google Colab

We have created a lexicon with 49 risk factors for suicidal thoughts and behaviors (plus one construct for kinship) validated by clinicians who are experts in suicide research.

fromconstruct_trackerimportlexiconsrl=lexicon.load_lexicon(name='srl_v1-0') # Load lexicondocuments= [
"I've been thinking about ending it all. I've been cutting. I just don't want to wake up.",
"I've been feeling all alone. No one cares about me. I've been hospitalized multiple times. I just want out. I'm pretty hopeless"
]
# Extractcounts, matches_by_construct, matches_doc2construct, matches_construct2doc=srl.extract(documents, normalize=False)
counts

Highlight matches

You can also access the Suicide Risk Lexicon in csv and json formats:


Create your own lexicon using generative AI

Open in Google Colab

Create a lexicon: keywords prototypically associated to a construct

We want to know if these documents contain mentions of certain construct "insight"

documents= [
"Every time I speak with my cousin Bob, I have great moments of clarity and wisdom", # mention of insight"He meditates a lot, but he's not super smart"# related to mindfulness, only somewhat related to insight"He is too competitive"] #not very related

Choose model here and obtain an API key from that provider. Cohere offers a free trial API key, 5 requests per minute. I'm going to choose GPT-4o:

os.environ["api_key"] ='YOUR_OPENAI_API_KEY'# This one might work for free models if no submissions have been tested: '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'model='gpt-4o'

Two lines of code to create a lexicon

l=lexicon.Lexicon() # Initialize lexiconl.add('Insight', section='tokens', value='create', source=model)

See results:

print(l.constructs['Insight']['tokens'])
['acuity', 'acumen', 'analysis', 'apprehension', 'awareness', 'clarity', 'comprehension', 'contemplation', 'depth', 'discernment', 'enlightenment', 'epiphany', 'foresight', 'grasp', 'illumination', 'insightfulness', 'interpretation', 'introspection', 'intuition', 'meditation', 'perception', 'perceptiveness', 'perspicacity', 'profoundness', 'realization', 'recognition', 'reflection', 'revelation', 'shrewdness', 'thoughtfulness', 'understanding', 'vision', 'wisdom']

We'll repeat for other constructs ("Mindfulness", "Compassion"). Now count whether tokens appear in document:

feature_vectors, matches_counter_d, matches_per_doc, matches_per_construct=lexicon.extract(
documents,
l.constructs,
normalize=False)
display(feature_vectors)

Lexicon counts

This traditional approach is perfectly interpretable. The first document contains three matches related to insight. Let's see which ones with highlight_matches():

lexicon.highlight_matches(documents, 'Insight', matches_construct2doc, max_matches=1)

Highlight matches



2. Construct-text similarity (CTS): finding similar phrases to tokens in your lexicon

Like Ctrl+F on steroids!

Lexicons may miss relevant words if not contained in the lexicon (it only counts exact matches). Embeddings can find semantically similar tokens. CTS will scan the document and return how similar is the most related phrase to any word in the lexicon.

magick -density 300 docs/images/cts.pdf -background white -alpha remove -quality 100 docs/images/cts.png Construct-text similarity

It will vectorize lexicon tokens and document tokens (e.g., phrases) into embeddings (quantitivae vector representing aspects of meaning). Then it will compute the similarity between both sets of tokens and return the maximum similarity as its score for the document.

lexicon_dict=my_lexicon.to_dict()
features, documents_tokenized, lexicon_dict_final_order, cosine_similarities=cts.measure(
lexicon_dict,
documents,
)
display(features)

Construct-text similarity

So we see that even though compassion did not find an exact match it had some relationship to the first two documents.

You can also sum the exact counts with the similarities for more fine-grained scores.

Construct-text similarity

We provide many features to add/remove tokens, generate definitions, validate with human ratings, and much more (see tutorials/construct_tracker.ipynb) Open in Google Colab


Structure of the lexicon.Lexicon() object

# Save general info on the lexiconmy_lexicon=lexicon.Lexicon() # Initialize lexiconmy_lexicon.name='Insight'# Set lexicon namemy_lexicon.description='Insight lexicon with constructs related to insight, mindfulness, and compassion'my_lexicon.creator='DML'# your name or initials for transparency in logging who made changesmy_lexicon.version='1.0'# Set version. Over time, others may modify your lexicon, so good to keep track. MAJOR.MINOR. (e.g., MAJOR: new constructs or big changes to a construct, Minor: small changes to a construct)# Each construct is a dict. You can save a lot of metadata depending on what you provide for each construct, for instance:print(my_lexicon.constructs)
{
'Insight': {
'variable_name': 'insight', # a name that is not sensitive to case with no spaces'prompt_name': 'insight',
'domain': 'psychology', # to guide Gen AI model as to sense of the construct (depression has different senses in psychology, geology, and economics)'examples': ['clarity', 'enlightenment', 'wise'], # to guide Gen AI model'definition': "the clarity of understanding of one's thoughts, feelings and behavior", # can be used in prompt and/or human validation'definition_references': 'Grant, A. M., Franklin, J., & Langford, P. (2002). The self-reflection and insight scale: A new measure of private self-consciousness. Social Behavior and Personality: an international journal, 30(8), 821-835.',
'tokens': ['acknowledgment',
'acuity',
'acumen',
'analytical',
'astute',
'awareness',
'clarity',
...],
'tokens_lemmatized': [], # when counting you can lemmatize all tokens for better results'remove': [], #which tokens to remove'tokens_metadata': {'gpt-4o-2024-05-13, temperature-0, ...': {
'action': 'create',
'tokens': [...],
'prompt': 'Provide many single words and some short phrases ...',
'time_elapsed': 14.21},
{'gpt-4o-2024-05-13, temperature-1, ...': { ... }},
}
},
'Mindfulness': {...},
'Compassion': {...},
}

Contributing

See docs/contributing.md

About

Track and measure constructs, concepts or categories in text documents

Resources

Contributing

Stars

11 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

BuildcodecovRuff

PyPIPython VersionLicense

construct-tracker

Track and measure constructs, concepts or categories in text documents. Build interpretable lexicon models quickly by using LLMs. Built on top of the OpenRouterAI package so you can use most Generative AI models.

Why build lexicons?

They can be used to build models that are:

  • interpretable: understand why the model outputs a given score, which can help avoid biases and guarantee the model will detect certain phrases (important for high-risk scenarios to use in tandem with LLMs)
  • lightweight: no GPU needed (unlike LLMs)
  • private and free: you can run on your local computer instead of submitting to a cloud API (OpenAI) which may not be secure
  • have high content validity: measure what you actually want to measure (unlike existing lexicons or models that measure something only slightly related)

If you use, please cite

Low DM, Rankin O, Coppersmith DDL, Bentley KH, Nock MK, Ghosh SS (2024). Using Generative AI to create lexicons for interpretable text models with high content validity. PsyarXiv.


Installation

pip install construct-tracker

Measure 49 suicide risk factors in text data

Highlight matches

TutorialOpen in Google Colab

We have created a lexicon with 49 risk factors for suicidal thoughts and behaviors (plus one construct for kinship) validated by clinicians who are experts in suicide research.

fromconstruct_trackerimportlexiconsrl=lexicon.load_lexicon(name='srl_v1-0') # Load lexicondocuments= [
"I've been thinking about ending it all. I've been cutting. I just don't want to wake up.",
"I've been feeling all alone. No one cares about me. I've been hospitalized multiple times. I just want out. I'm pretty hopeless"
]
# Extractcounts, matches_by_construct, matches_doc2construct, matches_construct2doc=srl.extract(documents, normalize=False)
counts

Highlight matches

You can also access the Suicide Risk Lexicon in csv and json formats:


Create your own lexicon using generative AI

Open in Google Colab

Create a lexicon: keywords prototypically associated to a construct

We want to know if these documents contain mentions of certain construct "insight"

documents= [
"Every time I speak with my cousin Bob, I have great moments of clarity and wisdom", # mention of insight"He meditates a lot, but he's not super smart"# related to mindfulness, only somewhat related to insight"He is too competitive"] #not very related

Choose model here and obtain an API key from that provider. Cohere offers a free trial API key, 5 requests per minute. I'm going to choose GPT-4o:

os.environ["api_key"] ='YOUR_OPENAI_API_KEY'# This one might work for free models if no submissions have been tested: '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'model='gpt-4o'

Two lines of code to create a lexicon

l=lexicon.Lexicon() # Initialize lexiconl.add('Insight', section='tokens', value='create', source=model)

See results:

print(l.constructs['Insight']['tokens'])
['acuity', 'acumen', 'analysis', 'apprehension', 'awareness', 'clarity', 'comprehension', 'contemplation', 'depth', 'discernment', 'enlightenment', 'epiphany', 'foresight', 'grasp', 'illumination', 'insightfulness', 'interpretation', 'introspection', 'intuition', 'meditation', 'perception', 'perceptiveness', 'perspicacity', 'profoundness', 'realization', 'recognition', 'reflection', 'revelation', 'shrewdness', 'thoughtfulness', 'understanding', 'vision', 'wisdom']

We'll repeat for other constructs ("Mindfulness", "Compassion"). Now count whether tokens appear in document:

feature_vectors, matches_counter_d, matches_per_doc, matches_per_construct=lexicon.extract(
documents,
l.constructs,
normalize=False)
display(feature_vectors)

Lexicon counts

This traditional approach is perfectly interpretable. The first document contains three matches related to insight. Let's see which ones with highlight_matches():

lexicon.highlight_matches(documents, 'Insight', matches_construct2doc, max_matches=1)

Highlight matches



2. Construct-text similarity (CTS): finding similar phrases to tokens in your lexicon

Like Ctrl+F on steroids!

Lexicons may miss relevant words if not contained in the lexicon (it only counts exact matches). Embeddings can find semantically similar tokens. CTS will scan the document and return how similar is the most related phrase to any word in the lexicon.

magick -density 300 docs/images/cts.pdf -background white -alpha remove -quality 100 docs/images/cts.png Construct-text similarity

It will vectorize lexicon tokens and document tokens (e.g., phrases) into embeddings (quantitivae vector representing aspects of meaning). Then it will compute the similarity between both sets of tokens and return the maximum similarity as its score for the document.

lexicon_dict=my_lexicon.to_dict()
features, documents_tokenized, lexicon_dict_final_order, cosine_similarities=cts.measure(
lexicon_dict,
documents,
)
display(features)

Construct-text similarity

So we see that even though compassion did not find an exact match it had some relationship to the first two documents.

You can also sum the exact counts with the similarities for more fine-grained scores.

Construct-text similarity

We provide many features to add/remove tokens, generate definitions, validate with human ratings, and much more (see tutorials/construct_tracker.ipynb) Open in Google Colab


Structure of the lexicon.Lexicon() object

# Save general info on the lexiconmy_lexicon=lexicon.Lexicon() # Initialize lexiconmy_lexicon.name='Insight'# Set lexicon namemy_lexicon.description='Insight lexicon with constructs related to insight, mindfulness, and compassion'my_lexicon.creator='DML'# your name or initials for transparency in logging who made changesmy_lexicon.version='1.0'# Set version. Over time, others may modify your lexicon, so good to keep track. MAJOR.MINOR. (e.g., MAJOR: new constructs or big changes to a construct, Minor: small changes to a construct)# Each construct is a dict. You can save a lot of metadata depending on what you provide for each construct, for instance:print(my_lexicon.constructs)
{
'Insight': {
'variable_name': 'insight', # a name that is not sensitive to case with no spaces'prompt_name': 'insight',
'domain': 'psychology', # to guide Gen AI model as to sense of the construct (depression has different senses in psychology, geology, and economics)'examples': ['clarity', 'enlightenment', 'wise'], # to guide Gen AI model'definition': "the clarity of understanding of one's thoughts, feelings and behavior", # can be used in prompt and/or human validation'definition_references': 'Grant, A. M., Franklin, J., & Langford, P. (2002). The self-reflection and insight scale: A new measure of private self-consciousness. Social Behavior and Personality: an international journal, 30(8), 821-835.',
'tokens': ['acknowledgment',
'acuity',
'acumen',
'analytical',
'astute',
'awareness',
'clarity',
...],
'tokens_lemmatized': [], # when counting you can lemmatize all tokens for better results'remove': [], #which tokens to remove'tokens_metadata': {'gpt-4o-2024-05-13, temperature-0, ...': {
'action': 'create',
'tokens': [...],
'prompt': 'Provide many single words and some short phrases ...',
'time_elapsed': 14.21},
{'gpt-4o-2024-05-13, temperature-1, ...': { ... }},
}
},
'Mindfulness': {...},
'Compassion': {...},
}

Contributing

See docs/contributing.md

About

Track and measure constructs, concepts or categories in text documents

Resources

Contributing

Stars

11 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

BuildcodecovRuff

PyPIPython VersionLicense

construct-tracker

Track and measure constructs, concepts or categories in text documents. Build interpretable lexicon models quickly by using LLMs. Built on top of the OpenRouterAI package so you can use most Generative AI models.

Why build lexicons?

They can be used to build models that are:

  • interpretable: understand why the model outputs a given score, which can help avoid biases and guarantee the model will detect certain phrases (important for high-risk scenarios to use in tandem with LLMs)
  • lightweight: no GPU needed (unlike LLMs)
  • private and free: you can run on your local computer instead of submitting to a cloud API (OpenAI) which may not be secure
  • have high content validity: measure what you actually want to measure (unlike existing lexicons or models that measure something only slightly related)

If you use, please cite

Low DM, Rankin O, Coppersmith DDL, Bentley KH, Nock MK, Ghosh SS (2024). Using Generative AI to create lexicons for interpretable text models with high content validity. PsyarXiv.


Installation

pip install construct-tracker

Measure 49 suicide risk factors in text data

Highlight matches

TutorialOpen in Google Colab

We have created a lexicon with 49 risk factors for suicidal thoughts and behaviors (plus one construct for kinship) validated by clinicians who are experts in suicide research.

fromconstruct_trackerimportlexiconsrl=lexicon.load_lexicon(name='srl_v1-0') # Load lexicondocuments= [
"I've been thinking about ending it all. I've been cutting. I just don't want to wake up.",
"I've been feeling all alone. No one cares about me. I've been hospitalized multiple times. I just want out. I'm pretty hopeless"
]
# Extractcounts, matches_by_construct, matches_doc2construct, matches_construct2doc=srl.extract(documents, normalize=False)
counts

Highlight matches

You can also access the Suicide Risk Lexicon in csv and json formats:


Create your own lexicon using generative AI

Open in Google Colab

Create a lexicon: keywords prototypically associated to a construct

We want to know if these documents contain mentions of certain construct "insight"

documents= [
"Every time I speak with my cousin Bob, I have great moments of clarity and wisdom", # mention of insight"He meditates a lot, but he's not super smart"# related to mindfulness, only somewhat related to insight"He is too competitive"] #not very related

Choose model here and obtain an API key from that provider. Cohere offers a free trial API key, 5 requests per minute. I'm going to choose GPT-4o:

os.environ["api_key"] ='YOUR_OPENAI_API_KEY'# This one might work for free models if no submissions have been tested: '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'model='gpt-4o'

Two lines of code to create a lexicon

l=lexicon.Lexicon() # Initialize lexiconl.add('Insight', section='tokens', value='create', source=model)

See results:

print(l.constructs['Insight']['tokens'])
['acuity', 'acumen', 'analysis', 'apprehension', 'awareness', 'clarity', 'comprehension', 'contemplation', 'depth', 'discernment', 'enlightenment', 'epiphany', 'foresight', 'grasp', 'illumination', 'insightfulness', 'interpretation', 'introspection', 'intuition', 'meditation', 'perception', 'perceptiveness', 'perspicacity', 'profoundness', 'realization', 'recognition', 'reflection', 'revelation', 'shrewdness', 'thoughtfulness', 'understanding', 'vision', 'wisdom']

We'll repeat for other constructs ("Mindfulness", "Compassion"). Now count whether tokens appear in document:

feature_vectors, matches_counter_d, matches_per_doc, matches_per_construct=lexicon.extract(
documents,
l.constructs,
normalize=False)
display(feature_vectors)

Lexicon counts

This traditional approach is perfectly interpretable. The first document contains three matches related to insight. Let's see which ones with highlight_matches():

lexicon.highlight_matches(documents, 'Insight', matches_construct2doc, max_matches=1)

Highlight matches



2. Construct-text similarity (CTS): finding similar phrases to tokens in your lexicon

Like Ctrl+F on steroids!

Lexicons may miss relevant words if not contained in the lexicon (it only counts exact matches). Embeddings can find semantically similar tokens. CTS will scan the document and return how similar is the most related phrase to any word in the lexicon.

magick -density 300 docs/images/cts.pdf -background white -alpha remove -quality 100 docs/images/cts.png Construct-text similarity

It will vectorize lexicon tokens and document tokens (e.g., phrases) into embeddings (quantitivae vector representing aspects of meaning). Then it will compute the similarity between both sets of tokens and return the maximum similarity as its score for the document.

lexicon_dict=my_lexicon.to_dict()
features, documents_tokenized, lexicon_dict_final_order, cosine_similarities=cts.measure(
lexicon_dict,
documents,
)
display(features)

Construct-text similarity

So we see that even though compassion did not find an exact match it had some relationship to the first two documents.

You can also sum the exact counts with the similarities for more fine-grained scores.

Construct-text similarity

We provide many features to add/remove tokens, generate definitions, validate with human ratings, and much more (see tutorials/construct_tracker.ipynb) Open in Google Colab


Structure of the lexicon.Lexicon() object

# Save general info on the lexiconmy_lexicon=lexicon.Lexicon() # Initialize lexiconmy_lexicon.name='Insight'# Set lexicon namemy_lexicon.description='Insight lexicon with constructs related to insight, mindfulness, and compassion'my_lexicon.creator='DML'# your name or initials for transparency in logging who made changesmy_lexicon.version='1.0'# Set version. Over time, others may modify your lexicon, so good to keep track. MAJOR.MINOR. (e.g., MAJOR: new constructs or big changes to a construct, Minor: small changes to a construct)# Each construct is a dict. You can save a lot of metadata depending on what you provide for each construct, for instance:print(my_lexicon.constructs)
{
'Insight': {
'variable_name': 'insight', # a name that is not sensitive to case with no spaces'prompt_name': 'insight',
'domain': 'psychology', # to guide Gen AI model as to sense of the construct (depression has different senses in psychology, geology, and economics)'examples': ['clarity', 'enlightenment', 'wise'], # to guide Gen AI model'definition': "the clarity of understanding of one's thoughts, feelings and behavior", # can be used in prompt and/or human validation'definition_references': 'Grant, A. M., Franklin, J., & Langford, P. (2002). The self-reflection and insight scale: A new measure of private self-consciousness. Social Behavior and Personality: an international journal, 30(8), 821-835.',
'tokens': ['acknowledgment',
'acuity',
'acumen',
'analytical',
'astute',
'awareness',
'clarity',
...],
'tokens_lemmatized': [], # when counting you can lemmatize all tokens for better results'remove': [], #which tokens to remove'tokens_metadata': {'gpt-4o-2024-05-13, temperature-0, ...': {
'action': 'create',
'tokens': [...],
'prompt': 'Provide many single words and some short phrases ...',
'time_elapsed': 14.21},
{'gpt-4o-2024-05-13, temperature-1, ...': { ... }},
}
},
'Mindfulness': {...},
'Compassion': {...},
}

Contributing

See docs/contributing.md

About

Track and measure constructs, concepts or categories in text documents

Resources

Contributing

Stars

11 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

BuildcodecovRuff

PyPIPython VersionLicense

construct-tracker

Track and measure constructs, concepts or categories in text documents. Build interpretable lexicon models quickly by using LLMs. Built on top of the OpenRouterAI package so you can use most Generative AI models.

Why build lexicons?

They can be used to build models that are:

  • interpretable: understand why the model outputs a given score, which can help avoid biases and guarantee the model will detect certain phrases (important for high-risk scenarios to use in tandem with LLMs)
  • lightweight: no GPU needed (unlike LLMs)
  • private and free: you can run on your local computer instead of submitting to a cloud API (OpenAI) which may not be secure
  • have high content validity: measure what you actually want to measure (unlike existing lexicons or models that measure something only slightly related)

If you use, please cite

Low DM, Rankin O, Coppersmith DDL, Bentley KH, Nock MK, Ghosh SS (2024). Using Generative AI to create lexicons for interpretable text models with high content validity. PsyarXiv.


Installation

pip install construct-tracker

Measure 49 suicide risk factors in text data

Highlight matches

TutorialOpen in Google Colab

We have created a lexicon with 49 risk factors for suicidal thoughts and behaviors (plus one construct for kinship) validated by clinicians who are experts in suicide research.

fromconstruct_trackerimportlexiconsrl=lexicon.load_lexicon(name='srl_v1-0') # Load lexicondocuments= [
"I've been thinking about ending it all. I've been cutting. I just don't want to wake up.",
"I've been feeling all alone. No one cares about me. I've been hospitalized multiple times. I just want out. I'm pretty hopeless"
]
# Extractcounts, matches_by_construct, matches_doc2construct, matches_construct2doc=srl.extract(documents, normalize=False)
counts

Highlight matches

You can also access the Suicide Risk Lexicon in csv and json formats:


Create your own lexicon using generative AI

Open in Google Colab

Create a lexicon: keywords prototypically associated to a construct

We want to know if these documents contain mentions of certain construct "insight"

documents= [
"Every time I speak with my cousin Bob, I have great moments of clarity and wisdom", # mention of insight"He meditates a lot, but he's not super smart"# related to mindfulness, only somewhat related to insight"He is too competitive"] #not very related

Choose model here and obtain an API key from that provider. Cohere offers a free trial API key, 5 requests per minute. I'm going to choose GPT-4o:

os.environ["api_key"] ='YOUR_OPENAI_API_KEY'# This one might work for free models if no submissions have been tested: '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'model='gpt-4o'

Two lines of code to create a lexicon

l=lexicon.Lexicon() # Initialize lexiconl.add('Insight', section='tokens', value='create', source=model)

See results:

print(l.constructs['Insight']['tokens'])
['acuity', 'acumen', 'analysis', 'apprehension', 'awareness', 'clarity', 'comprehension', 'contemplation', 'depth', 'discernment', 'enlightenment', 'epiphany', 'foresight', 'grasp', 'illumination', 'insightfulness', 'interpretation', 'introspection', 'intuition', 'meditation', 'perception', 'perceptiveness', 'perspicacity', 'profoundness', 'realization', 'recognition', 'reflection', 'revelation', 'shrewdness', 'thoughtfulness', 'understanding', 'vision', 'wisdom']

We'll repeat for other constructs ("Mindfulness", "Compassion"). Now count whether tokens appear in document:

feature_vectors, matches_counter_d, matches_per_doc, matches_per_construct=lexicon.extract(
documents,
l.constructs,
normalize=False)
display(feature_vectors)

Lexicon counts

This traditional approach is perfectly interpretable. The first document contains three matches related to insight. Let's see which ones with highlight_matches():

lexicon.highlight_matches(documents, 'Insight', matches_construct2doc, max_matches=1)

Highlight matches



2. Construct-text similarity (CTS): finding similar phrases to tokens in your lexicon

Like Ctrl+F on steroids!

Lexicons may miss relevant words if not contained in the lexicon (it only counts exact matches). Embeddings can find semantically similar tokens. CTS will scan the document and return how similar is the most related phrase to any word in the lexicon.

magick -density 300 docs/images/cts.pdf -background white -alpha remove -quality 100 docs/images/cts.png Construct-text similarity

It will vectorize lexicon tokens and document tokens (e.g., phrases) into embeddings (quantitivae vector representing aspects of meaning). Then it will compute the similarity between both sets of tokens and return the maximum similarity as its score for the document.

lexicon_dict=my_lexicon.to_dict()
features, documents_tokenized, lexicon_dict_final_order, cosine_similarities=cts.measure(
lexicon_dict,
documents,
)
display(features)

Construct-text similarity

So we see that even though compassion did not find an exact match it had some relationship to the first two documents.

You can also sum the exact counts with the similarities for more fine-grained scores.

Construct-text similarity

We provide many features to add/remove tokens, generate definitions, validate with human ratings, and much more (see tutorials/construct_tracker.ipynb) Open in Google Colab


Structure of the lexicon.Lexicon() object

# Save general info on the lexiconmy_lexicon=lexicon.Lexicon() # Initialize lexiconmy_lexicon.name='Insight'# Set lexicon namemy_lexicon.description='Insight lexicon with constructs related to insight, mindfulness, and compassion'my_lexicon.creator='DML'# your name or initials for transparency in logging who made changesmy_lexicon.version='1.0'# Set version. Over time, others may modify your lexicon, so good to keep track. MAJOR.MINOR. (e.g., MAJOR: new constructs or big changes to a construct, Minor: small changes to a construct)# Each construct is a dict. You can save a lot of metadata depending on what you provide for each construct, for instance:print(my_lexicon.constructs)
{
'Insight': {
'variable_name': 'insight', # a name that is not sensitive to case with no spaces'prompt_name': 'insight',
'domain': 'psychology', # to guide Gen AI model as to sense of the construct (depression has different senses in psychology, geology, and economics)'examples': ['clarity', 'enlightenment', 'wise'], # to guide Gen AI model'definition': "the clarity of understanding of one's thoughts, feelings and behavior", # can be used in prompt and/or human validation'definition_references': 'Grant, A. M., Franklin, J., & Langford, P. (2002). The self-reflection and insight scale: A new measure of private self-consciousness. Social Behavior and Personality: an international journal, 30(8), 821-835.',
'tokens': ['acknowledgment',
'acuity',
'acumen',
'analytical',
'astute',
'awareness',
'clarity',
...],
'tokens_lemmatized': [], # when counting you can lemmatize all tokens for better results'remove': [], #which tokens to remove'tokens_metadata': {'gpt-4o-2024-05-13, temperature-0, ...': {
'action': 'create',
'tokens': [...],
'prompt': 'Provide many single words and some short phrases ...',
'time_elapsed': 14.21},
{'gpt-4o-2024-05-13, temperature-1, ...': { ... }},
}
},
'Mindfulness': {...},
'Compassion': {...},
}

Contributing

See docs/contributing.md

About

Track and measure constructs, concepts or categories in text documents

Resources

Contributing

Stars

11 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

BuildcodecovRuff

PyPIPython VersionLicense

construct-tracker

Track and measure constructs, concepts or categories in text documents. Build interpretable lexicon models quickly by using LLMs. Built on top of the OpenRouterAI package so you can use most Generative AI models.

Why build lexicons?

They can be used to build models that are:

  • interpretable: understand why the model outputs a given score, which can help avoid biases and guarantee the model will detect certain phrases (important for high-risk scenarios to use in tandem with LLMs)
  • lightweight: no GPU needed (unlike LLMs)
  • private and free: you can run on your local computer instead of submitting to a cloud API (OpenAI) which may not be secure
  • have high content validity: measure what you actually want to measure (unlike existing lexicons or models that measure something only slightly related)

If you use, please cite

Low DM, Rankin O, Coppersmith DDL, Bentley KH, Nock MK, Ghosh SS (2024). Using Generative AI to create lexicons for interpretable text models with high content validity. PsyarXiv.


Installation

pip install construct-tracker

Measure 49 suicide risk factors in text data

Highlight matches

TutorialOpen in Google Colab

We have created a lexicon with 49 risk factors for suicidal thoughts and behaviors (plus one construct for kinship) validated by clinicians who are experts in suicide research.

fromconstruct_trackerimportlexiconsrl=lexicon.load_lexicon(name='srl_v1-0') # Load lexicondocuments= [
"I've been thinking about ending it all. I've been cutting. I just don't want to wake up.",
"I've been feeling all alone. No one cares about me. I've been hospitalized multiple times. I just want out. I'm pretty hopeless"
]
# Extractcounts, matches_by_construct, matches_doc2construct, matches_construct2doc=srl.extract(documents, normalize=False)
counts

Highlight matches

You can also access the Suicide Risk Lexicon in csv and json formats:


Create your own lexicon using generative AI

Open in Google Colab

Create a lexicon: keywords prototypically associated to a construct

We want to know if these documents contain mentions of certain construct "insight"

documents= [
"Every time I speak with my cousin Bob, I have great moments of clarity and wisdom", # mention of insight"He meditates a lot, but he's not super smart"# related to mindfulness, only somewhat related to insight"He is too competitive"] #not very related

Choose model here and obtain an API key from that provider. Cohere offers a free trial API key, 5 requests per minute. I'm going to choose GPT-4o:

os.environ["api_key"] ='YOUR_OPENAI_API_KEY'# This one might work for free models if no submissions have been tested: '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'model='gpt-4o'

Two lines of code to create a lexicon

l=lexicon.Lexicon() # Initialize lexiconl.add('Insight', section='tokens', value='create', source=model)

See results:

print(l.constructs['Insight']['tokens'])
['acuity', 'acumen', 'analysis', 'apprehension', 'awareness', 'clarity', 'comprehension', 'contemplation', 'depth', 'discernment', 'enlightenment', 'epiphany', 'foresight', 'grasp', 'illumination', 'insightfulness', 'interpretation', 'introspection', 'intuition', 'meditation', 'perception', 'perceptiveness', 'perspicacity', 'profoundness', 'realization', 'recognition', 'reflection', 'revelation', 'shrewdness', 'thoughtfulness', 'understanding', 'vision', 'wisdom']

We'll repeat for other constructs ("Mindfulness", "Compassion"). Now count whether tokens appear in document:

feature_vectors, matches_counter_d, matches_per_doc, matches_per_construct=lexicon.extract(
documents,
l.constructs,
normalize=False)
display(feature_vectors)

Lexicon counts

This traditional approach is perfectly interpretable. The first document contains three matches related to insight. Let's see which ones with highlight_matches():

lexicon.highlight_matches(documents, 'Insight', matches_construct2doc, max_matches=1)

Highlight matches



2. Construct-text similarity (CTS): finding similar phrases to tokens in your lexicon

Like Ctrl+F on steroids!

Lexicons may miss relevant words if not contained in the lexicon (it only counts exact matches). Embeddings can find semantically similar tokens. CTS will scan the document and return how similar is the most related phrase to any word in the lexicon.

magick -density 300 docs/images/cts.pdf -background white -alpha remove -quality 100 docs/images/cts.png Construct-text similarity

It will vectorize lexicon tokens and document tokens (e.g., phrases) into embeddings (quantitivae vector representing aspects of meaning). Then it will compute the similarity between both sets of tokens and return the maximum similarity as its score for the document.

lexicon_dict=my_lexicon.to_dict()
features, documents_tokenized, lexicon_dict_final_order, cosine_similarities=cts.measure(
lexicon_dict,
documents,
)
display(features)

Construct-text similarity

So we see that even though compassion did not find an exact match it had some relationship to the first two documents.

You can also sum the exact counts with the similarities for more fine-grained scores.

Construct-text similarity

We provide many features to add/remove tokens, generate definitions, validate with human ratings, and much more (see tutorials/construct_tracker.ipynb) Open in Google Colab


Structure of the lexicon.Lexicon() object

# Save general info on the lexiconmy_lexicon=lexicon.Lexicon() # Initialize lexiconmy_lexicon.name='Insight'# Set lexicon namemy_lexicon.description='Insight lexicon with constructs related to insight, mindfulness, and compassion'my_lexicon.creator='DML'# your name or initials for transparency in logging who made changesmy_lexicon.version='1.0'# Set version. Over time, others may modify your lexicon, so good to keep track. MAJOR.MINOR. (e.g., MAJOR: new constructs or big changes to a construct, Minor: small changes to a construct)# Each construct is a dict. You can save a lot of metadata depending on what you provide for each construct, for instance:print(my_lexicon.constructs)
{
'Insight': {
'variable_name': 'insight', # a name that is not sensitive to case with no spaces'prompt_name': 'insight',
'domain': 'psychology', # to guide Gen AI model as to sense of the construct (depression has different senses in psychology, geology, and economics)'examples': ['clarity', 'enlightenment', 'wise'], # to guide Gen AI model'definition': "the clarity of understanding of one's thoughts, feelings and behavior", # can be used in prompt and/or human validation'definition_references': 'Grant, A. M., Franklin, J., & Langford, P. (2002). The self-reflection and insight scale: A new measure of private self-consciousness. Social Behavior and Personality: an international journal, 30(8), 821-835.',
'tokens': ['acknowledgment',
'acuity',
'acumen',
'analytical',
'astute',
'awareness',
'clarity',
...],
'tokens_lemmatized': [], # when counting you can lemmatize all tokens for better results'remove': [], #which tokens to remove'tokens_metadata': {'gpt-4o-2024-05-13, temperature-0, ...': {
'action': 'create',
'tokens': [...],
'prompt': 'Provide many single words and some short phrases ...',
'time_elapsed': 14.21},
{'gpt-4o-2024-05-13, temperature-1, ...': { ... }},
}
},
'Mindfulness': {...},
'Compassion': {...},
}

Contributing

See docs/contributing.md

About

Track and measure constructs, concepts or categories in text documents

Resources

Contributing

Stars

11 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

BuildcodecovRuff

PyPIPython VersionLicense

construct-tracker

Track and measure constructs, concepts or categories in text documents. Build interpretable lexicon models quickly by using LLMs. Built on top of the OpenRouterAI package so you can use most Generative AI models.

Why build lexicons?

They can be used to build models that are:

  • interpretable: understand why the model outputs a given score, which can help avoid biases and guarantee the model will detect certain phrases (important for high-risk scenarios to use in tandem with LLMs)
  • lightweight: no GPU needed (unlike LLMs)
  • private and free: you can run on your local computer instead of submitting to a cloud API (OpenAI) which may not be secure
  • have high content validity: measure what you actually want to measure (unlike existing lexicons or models that measure something only slightly related)

If you use, please cite

Low DM, Rankin O, Coppersmith DDL, Bentley KH, Nock MK, Ghosh SS (2024). Using Generative AI to create lexicons for interpretable text models with high content validity. PsyarXiv.


Installation

pip install construct-tracker

Measure 49 suicide risk factors in text data

Highlight matches

TutorialOpen in Google Colab

We have created a lexicon with 49 risk factors for suicidal thoughts and behaviors (plus one construct for kinship) validated by clinicians who are experts in suicide research.

fromconstruct_trackerimportlexiconsrl=lexicon.load_lexicon(name='srl_v1-0') # Load lexicondocuments= [
"I've been thinking about ending it all. I've been cutting. I just don't want to wake up.",
"I've been feeling all alone. No one cares about me. I've been hospitalized multiple times. I just want out. I'm pretty hopeless"
]
# Extractcounts, matches_by_construct, matches_doc2construct, matches_construct2doc=srl.extract(documents, normalize=False)
counts

Highlight matches

You can also access the Suicide Risk Lexicon in csv and json formats:


Create your own lexicon using generative AI

Open in Google Colab

Create a lexicon: keywords prototypically associated to a construct

We want to know if these documents contain mentions of certain construct "insight"

documents= [
"Every time I speak with my cousin Bob, I have great moments of clarity and wisdom", # mention of insight"He meditates a lot, but he's not super smart"# related to mindfulness, only somewhat related to insight"He is too competitive"] #not very related

Choose model here and obtain an API key from that provider. Cohere offers a free trial API key, 5 requests per minute. I'm going to choose GPT-4o:

os.environ["api_key"] ='YOUR_OPENAI_API_KEY'# This one might work for free models if no submissions have been tested: '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'model='gpt-4o'

Two lines of code to create a lexicon

l=lexicon.Lexicon() # Initialize lexiconl.add('Insight', section='tokens', value='create', source=model)

See results:

print(l.constructs['Insight']['tokens'])
['acuity', 'acumen', 'analysis', 'apprehension', 'awareness', 'clarity', 'comprehension', 'contemplation', 'depth', 'discernment', 'enlightenment', 'epiphany', 'foresight', 'grasp', 'illumination', 'insightfulness', 'interpretation', 'introspection', 'intuition', 'meditation', 'perception', 'perceptiveness', 'perspicacity', 'profoundness', 'realization', 'recognition', 'reflection', 'revelation', 'shrewdness', 'thoughtfulness', 'understanding', 'vision', 'wisdom']

We'll repeat for other constructs ("Mindfulness", "Compassion"). Now count whether tokens appear in document:

feature_vectors, matches_counter_d, matches_per_doc, matches_per_construct=lexicon.extract(
documents,
l.constructs,
normalize=False)
display(feature_vectors)

Lexicon counts

This traditional approach is perfectly interpretable. The first document contains three matches related to insight. Let's see which ones with highlight_matches():

lexicon.highlight_matches(documents, 'Insight', matches_construct2doc, max_matches=1)

Highlight matches



2. Construct-text similarity (CTS): finding similar phrases to tokens in your lexicon

Like Ctrl+F on steroids!

Lexicons may miss relevant words if not contained in the lexicon (it only counts exact matches). Embeddings can find semantically similar tokens. CTS will scan the document and return how similar is the most related phrase to any word in the lexicon.

magick -density 300 docs/images/cts.pdf -background white -alpha remove -quality 100 docs/images/cts.png Construct-text similarity

It will vectorize lexicon tokens and document tokens (e.g., phrases) into embeddings (quantitivae vector representing aspects of meaning). Then it will compute the similarity between both sets of tokens and return the maximum similarity as its score for the document.

lexicon_dict=my_lexicon.to_dict()
features, documents_tokenized, lexicon_dict_final_order, cosine_similarities=cts.measure(
lexicon_dict,
documents,
)
display(features)

Construct-text similarity

So we see that even though compassion did not find an exact match it had some relationship to the first two documents.

You can also sum the exact counts with the similarities for more fine-grained scores.

Construct-text similarity

We provide many features to add/remove tokens, generate definitions, validate with human ratings, and much more (see tutorials/construct_tracker.ipynb) Open in Google Colab


Structure of the lexicon.Lexicon() object

# Save general info on the lexiconmy_lexicon=lexicon.Lexicon() # Initialize lexiconmy_lexicon.name='Insight'# Set lexicon namemy_lexicon.description='Insight lexicon with constructs related to insight, mindfulness, and compassion'my_lexicon.creator='DML'# your name or initials for transparency in logging who made changesmy_lexicon.version='1.0'# Set version. Over time, others may modify your lexicon, so good to keep track. MAJOR.MINOR. (e.g., MAJOR: new constructs or big changes to a construct, Minor: small changes to a construct)# Each construct is a dict. You can save a lot of metadata depending on what you provide for each construct, for instance:print(my_lexicon.constructs)
{
'Insight': {
'variable_name': 'insight', # a name that is not sensitive to case with no spaces'prompt_name': 'insight',
'domain': 'psychology', # to guide Gen AI model as to sense of the construct (depression has different senses in psychology, geology, and economics)'examples': ['clarity', 'enlightenment', 'wise'], # to guide Gen AI model'definition': "the clarity of understanding of one's thoughts, feelings and behavior", # can be used in prompt and/or human validation'definition_references': 'Grant, A. M., Franklin, J., & Langford, P. (2002). The self-reflection and insight scale: A new measure of private self-consciousness. Social Behavior and Personality: an international journal, 30(8), 821-835.',
'tokens': ['acknowledgment',
'acuity',
'acumen',
'analytical',
'astute',
'awareness',
'clarity',
...],
'tokens_lemmatized': [], # when counting you can lemmatize all tokens for better results'remove': [], #which tokens to remove'tokens_metadata': {'gpt-4o-2024-05-13, temperature-0, ...': {
'action': 'create',
'tokens': [...],
'prompt': 'Provide many single words and some short phrases ...',
'time_elapsed': 14.21},
{'gpt-4o-2024-05-13, temperature-1, ...': { ... }},
}
},
'Mindfulness': {...},
'Compassion': {...},
}

Contributing

See docs/contributing.md

About

Track and measure constructs, concepts or categories in text documents

Resources

Contributing

Stars

11 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

BuildcodecovRuff

PyPIPython VersionLicense

construct-tracker

Track and measure constructs, concepts or categories in text documents. Build interpretable lexicon models quickly by using LLMs. Built on top of the OpenRouterAI package so you can use most Generative AI models.

Why build lexicons?

They can be used to build models that are:

  • interpretable: understand why the model outputs a given score, which can help avoid biases and guarantee the model will detect certain phrases (important for high-risk scenarios to use in tandem with LLMs)
  • lightweight: no GPU needed (unlike LLMs)
  • private and free: you can run on your local computer instead of submitting to a cloud API (OpenAI) which may not be secure
  • have high content validity: measure what you actually want to measure (unlike existing lexicons or models that measure something only slightly related)

If you use, please cite

Low DM, Rankin O, Coppersmith DDL, Bentley KH, Nock MK, Ghosh SS (2024). Using Generative AI to create lexicons for interpretable text models with high content validity. PsyarXiv.


Installation

pip install construct-tracker

Measure 49 suicide risk factors in text data

Highlight matches

TutorialOpen in Google Colab

We have created a lexicon with 49 risk factors for suicidal thoughts and behaviors (plus one construct for kinship) validated by clinicians who are experts in suicide research.

fromconstruct_trackerimportlexiconsrl=lexicon.load_lexicon(name='srl_v1-0') # Load lexicondocuments= [
"I've been thinking about ending it all. I've been cutting. I just don't want to wake up.",
"I've been feeling all alone. No one cares about me. I've been hospitalized multiple times. I just want out. I'm pretty hopeless"
]
# Extractcounts, matches_by_construct, matches_doc2construct, matches_construct2doc=srl.extract(documents, normalize=False)
counts

Highlight matches

You can also access the Suicide Risk Lexicon in csv and json formats:


Create your own lexicon using generative AI

Open in Google Colab

Create a lexicon: keywords prototypically associated to a construct

We want to know if these documents contain mentions of certain construct "insight"

documents= [
"Every time I speak with my cousin Bob, I have great moments of clarity and wisdom", # mention of insight"He meditates a lot, but he's not super smart"# related to mindfulness, only somewhat related to insight"He is too competitive"] #not very related

Choose model here and obtain an API key from that provider. Cohere offers a free trial API key, 5 requests per minute. I'm going to choose GPT-4o:

os.environ["api_key"] ='YOUR_OPENAI_API_KEY'# This one might work for free models if no submissions have been tested: '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'model='gpt-4o'

Two lines of code to create a lexicon

l=lexicon.Lexicon() # Initialize lexiconl.add('Insight', section='tokens', value='create', source=model)

See results:

print(l.constructs['Insight']['tokens'])
['acuity', 'acumen', 'analysis', 'apprehension', 'awareness', 'clarity', 'comprehension', 'contemplation', 'depth', 'discernment', 'enlightenment', 'epiphany', 'foresight', 'grasp', 'illumination', 'insightfulness', 'interpretation', 'introspection', 'intuition', 'meditation', 'perception', 'perceptiveness', 'perspicacity', 'profoundness', 'realization', 'recognition', 'reflection', 'revelation', 'shrewdness', 'thoughtfulness', 'understanding', 'vision', 'wisdom']

We'll repeat for other constructs ("Mindfulness", "Compassion"). Now count whether tokens appear in document:

feature_vectors, matches_counter_d, matches_per_doc, matches_per_construct=lexicon.extract(
documents,
l.constructs,
normalize=False)
display(feature_vectors)

Lexicon counts

This traditional approach is perfectly interpretable. The first document contains three matches related to insight. Let's see which ones with highlight_matches():

lexicon.highlight_matches(documents, 'Insight', matches_construct2doc, max_matches=1)

Highlight matches



2. Construct-text similarity (CTS): finding similar phrases to tokens in your lexicon

Like Ctrl+F on steroids!

Lexicons may miss relevant words if not contained in the lexicon (it only counts exact matches). Embeddings can find semantically similar tokens. CTS will scan the document and return how similar is the most related phrase to any word in the lexicon.

magick -density 300 docs/images/cts.pdf -background white -alpha remove -quality 100 docs/images/cts.png Construct-text similarity

It will vectorize lexicon tokens and document tokens (e.g., phrases) into embeddings (quantitivae vector representing aspects of meaning). Then it will compute the similarity between both sets of tokens and return the maximum similarity as its score for the document.

lexicon_dict=my_lexicon.to_dict()
features, documents_tokenized, lexicon_dict_final_order, cosine_similarities=cts.measure(
lexicon_dict,
documents,
)
display(features)

Construct-text similarity

So we see that even though compassion did not find an exact match it had some relationship to the first two documents.

You can also sum the exact counts with the similarities for more fine-grained scores.

Construct-text similarity

We provide many features to add/remove tokens, generate definitions, validate with human ratings, and much more (see tutorials/construct_tracker.ipynb) Open in Google Colab


Structure of the lexicon.Lexicon() object

# Save general info on the lexiconmy_lexicon=lexicon.Lexicon() # Initialize lexiconmy_lexicon.name='Insight'# Set lexicon namemy_lexicon.description='Insight lexicon with constructs related to insight, mindfulness, and compassion'my_lexicon.creator='DML'# your name or initials for transparency in logging who made changesmy_lexicon.version='1.0'# Set version. Over time, others may modify your lexicon, so good to keep track. MAJOR.MINOR. (e.g., MAJOR: new constructs or big changes to a construct, Minor: small changes to a construct)# Each construct is a dict. You can save a lot of metadata depending on what you provide for each construct, for instance:print(my_lexicon.constructs)
{
'Insight': {
'variable_name': 'insight', # a name that is not sensitive to case with no spaces'prompt_name': 'insight',
'domain': 'psychology', # to guide Gen AI model as to sense of the construct (depression has different senses in psychology, geology, and economics)'examples': ['clarity', 'enlightenment', 'wise'], # to guide Gen AI model'definition': "the clarity of understanding of one's thoughts, feelings and behavior", # can be used in prompt and/or human validation'definition_references': 'Grant, A. M., Franklin, J., & Langford, P. (2002). The self-reflection and insight scale: A new measure of private self-consciousness. Social Behavior and Personality: an international journal, 30(8), 821-835.',
'tokens': ['acknowledgment',
'acuity',
'acumen',
'analytical',
'astute',
'awareness',
'clarity',
...],
'tokens_lemmatized': [], # when counting you can lemmatize all tokens for better results'remove': [], #which tokens to remove'tokens_metadata': {'gpt-4o-2024-05-13, temperature-0, ...': {
'action': 'create',
'tokens': [...],
'prompt': 'Provide many single words and some short phrases ...',
'time_elapsed': 14.21},
{'gpt-4o-2024-05-13, temperature-1, ...': { ... }},
}
},
'Mindfulness': {...},
'Compassion': {...},
}

Contributing

See docs/contributing.md

About

Track and measure constructs, concepts or categories in text documents

Resources

Contributing

Stars

11 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

BuildcodecovRuff

PyPIPython VersionLicense

construct-tracker

Track and measure constructs, concepts or categories in text documents. Build interpretable lexicon models quickly by using LLMs. Built on top of the OpenRouterAI package so you can use most Generative AI models.

Why build lexicons?

They can be used to build models that are:

  • interpretable: understand why the model outputs a given score, which can help avoid biases and guarantee the model will detect certain phrases (important for high-risk scenarios to use in tandem with LLMs)
  • lightweight: no GPU needed (unlike LLMs)
  • private and free: you can run on your local computer instead of submitting to a cloud API (OpenAI) which may not be secure
  • have high content validity: measure what you actually want to measure (unlike existing lexicons or models that measure something only slightly related)

If you use, please cite

Low DM, Rankin O, Coppersmith DDL, Bentley KH, Nock MK, Ghosh SS (2024). Using Generative AI to create lexicons for interpretable text models with high content validity. PsyarXiv.


Installation

pip install construct-tracker

Measure 49 suicide risk factors in text data

Highlight matches

TutorialOpen in Google Colab

We have created a lexicon with 49 risk factors for suicidal thoughts and behaviors (plus one construct for kinship) validated by clinicians who are experts in suicide research.

fromconstruct_trackerimportlexiconsrl=lexicon.load_lexicon(name='srl_v1-0') # Load lexicondocuments= [
"I've been thinking about ending it all. I've been cutting. I just don't want to wake up.",
"I've been feeling all alone. No one cares about me. I've been hospitalized multiple times. I just want out. I'm pretty hopeless"
]
# Extractcounts, matches_by_construct, matches_doc2construct, matches_construct2doc=srl.extract(documents, normalize=False)
counts

Highlight matches

You can also access the Suicide Risk Lexicon in csv and json formats:


Create your own lexicon using generative AI

Open in Google Colab

Create a lexicon: keywords prototypically associated to a construct

We want to know if these documents contain mentions of certain construct "insight"

documents= [
"Every time I speak with my cousin Bob, I have great moments of clarity and wisdom", # mention of insight"He meditates a lot, but he's not super smart"# related to mindfulness, only somewhat related to insight"He is too competitive"] #not very related

Choose model here and obtain an API key from that provider. Cohere offers a free trial API key, 5 requests per minute. I'm going to choose GPT-4o:

os.environ["api_key"] ='YOUR_OPENAI_API_KEY'# This one might work for free models if no submissions have been tested: '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'model='gpt-4o'

Two lines of code to create a lexicon

l=lexicon.Lexicon() # Initialize lexiconl.add('Insight', section='tokens', value='create', source=model)

See results:

print(l.constructs['Insight']['tokens'])
['acuity', 'acumen', 'analysis', 'apprehension', 'awareness', 'clarity', 'comprehension', 'contemplation', 'depth', 'discernment', 'enlightenment', 'epiphany', 'foresight', 'grasp', 'illumination', 'insightfulness', 'interpretation', 'introspection', 'intuition', 'meditation', 'perception', 'perceptiveness', 'perspicacity', 'profoundness', 'realization', 'recognition', 'reflection', 'revelation', 'shrewdness', 'thoughtfulness', 'understanding', 'vision', 'wisdom']

We'll repeat for other constructs ("Mindfulness", "Compassion"). Now count whether tokens appear in document:

feature_vectors, matches_counter_d, matches_per_doc, matches_per_construct=lexicon.extract(
documents,
l.constructs,
normalize=False)
display(feature_vectors)

Lexicon counts

This traditional approach is perfectly interpretable. The first document contains three matches related to insight. Let's see which ones with highlight_matches():

lexicon.highlight_matches(documents, 'Insight', matches_construct2doc, max_matches=1)

Highlight matches



2. Construct-text similarity (CTS): finding similar phrases to tokens in your lexicon

Like Ctrl+F on steroids!

Lexicons may miss relevant words if not contained in the lexicon (it only counts exact matches). Embeddings can find semantically similar tokens. CTS will scan the document and return how similar is the most related phrase to any word in the lexicon.

magick -density 300 docs/images/cts.pdf -background white -alpha remove -quality 100 docs/images/cts.png Construct-text similarity

It will vectorize lexicon tokens and document tokens (e.g., phrases) into embeddings (quantitivae vector representing aspects of meaning). Then it will compute the similarity between both sets of tokens and return the maximum similarity as its score for the document.

lexicon_dict=my_lexicon.to_dict()
features, documents_tokenized, lexicon_dict_final_order, cosine_similarities=cts.measure(
lexicon_dict,
documents,
)
display(features)

Construct-text similarity

So we see that even though compassion did not find an exact match it had some relationship to the first two documents.

You can also sum the exact counts with the similarities for more fine-grained scores.

Construct-text similarity

We provide many features to add/remove tokens, generate definitions, validate with human ratings, and much more (see tutorials/construct_tracker.ipynb) Open in Google Colab


Structure of the lexicon.Lexicon() object

# Save general info on the lexiconmy_lexicon=lexicon.Lexicon() # Initialize lexiconmy_lexicon.name='Insight'# Set lexicon namemy_lexicon.description='Insight lexicon with constructs related to insight, mindfulness, and compassion'my_lexicon.creator='DML'# your name or initials for transparency in logging who made changesmy_lexicon.version='1.0'# Set version. Over time, others may modify your lexicon, so good to keep track. MAJOR.MINOR. (e.g., MAJOR: new constructs or big changes to a construct, Minor: small changes to a construct)# Each construct is a dict. You can save a lot of metadata depending on what you provide for each construct, for instance:print(my_lexicon.constructs)
{
'Insight': {
'variable_name': 'insight', # a name that is not sensitive to case with no spaces'prompt_name': 'insight',
'domain': 'psychology', # to guide Gen AI model as to sense of the construct (depression has different senses in psychology, geology, and economics)'examples': ['clarity', 'enlightenment', 'wise'], # to guide Gen AI model'definition': "the clarity of understanding of one's thoughts, feelings and behavior", # can be used in prompt and/or human validation'definition_references': 'Grant, A. M., Franklin, J., & Langford, P. (2002). The self-reflection and insight scale: A new measure of private self-consciousness. Social Behavior and Personality: an international journal, 30(8), 821-835.',
'tokens': ['acknowledgment',
'acuity',
'acumen',
'analytical',
'astute',
'awareness',
'clarity',
...],
'tokens_lemmatized': [], # when counting you can lemmatize all tokens for better results'remove': [], #which tokens to remove'tokens_metadata': {'gpt-4o-2024-05-13, temperature-0, ...': {
'action': 'create',
'tokens': [...],
'prompt': 'Provide many single words and some short phrases ...',
'time_elapsed': 14.21},
{'gpt-4o-2024-05-13, temperature-1, ...': { ... }},
}
},
'Mindfulness': {...},
'Compassion': {...},
}

Contributing

See docs/contributing.md

About

Track and measure constructs, concepts or categories in text documents

Resources

Contributing

Stars

11 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages