Repository files navigation

Feel free to use, but this is in beta version and is still being tested.

LLM Tracker

LLM Tracker is a Python package for identifying psychological constructs in text data (e.g., interviews, social media posts, chatbot interactions) and comparing LLM-coded results against human-coded results.

The package supports:

  • Access to thousands of remote or local LLMs through Mozilla's any-llm API (OpenRouter, Azure, Ollama, etc.)
  • Detect every instance of a construct (e.g., anxiety) and return the verbatim quote (e.g., "I'm worried about my cousin")
  • Comparing human and LLM codings at the quote level (using LLMs to match human and LLM quotes)
  • Computing inter-rater reliability metrics (Kappa, ICC, PABAK) and classification metrics (sensitivity, precision, F1, and PR AUC) and returning summary tables
  • Automatically retrying submissions when LLM outputs are not parseable
  • Saving analyzer outputs, metadata, and retryable error records
  • Flexible loading of csv, txt, docx and preprocessing of dedoose human coding to match LLM coding dataframes.
  • Many new features coming soon: visualizations, automated prompt engineering, and more!

Please cite this if you use this package:

Low, D., Mair, P., Nock, M., & Ghosh, S. (2025). Text Psychometrics: Assessing Psychological Constructs in Text Using Natural Language Processing. PsyArxiv. https://osf.io/preprints/psyarxiv/9rdux_v4

Installation

Install dependencies with Poetry:

poetry install

For tutorial extras such as corpus summaries:

poetry install --with tutorials

API Key

LLM Tracker uses any-llm API for LLM calls, which can run the most common APIs (OpenRouter, Azure, Ollama). For instance, you can add an OpenRouter API key by adding a few dollars here https://openrouter.ai/. Each LLM call tends to cost a a fraction of a cent (see cost for specific models on OpenRouter). Provide an API key directly:

api_key="your-openrouter-key"

or set it in the environment:

export OPENROUTER_API_KEY="your-openrouter-key"

You can also pass the path to a .env file containing:

OPENROUTER_API_KEY=your-openrouter-key

Basic LLM Coding

fromllm_trackerimportLLMTrackerAnalyzeranalyzer=LLMTrackerAnalyzer(
api_key=api_key,
model_name="google/gemini-3-flash-preview",
)
results_llm, metadata_llm, errors_llm=analyzer.analyze_csv(
csv_path="sample_data.csv",
codebook_path="codebook.json",
text_column="post",
subreddit_column="subreddit",
author_column="author",
output_dir="LLM_coding",
)

For a directory of supported document files:

results_llm, metadata_llm, errors_llm=analyzer.analyze_directory(
input_dir="documents",
codebook_path="codebook.json",
output_dir="LLM_coding",
)

Directory mode supports .txt and .csv files. Each file becomes one document.

Human Coding Input

Human coding is loaded into memory and passed directly to the comparer:

fromllm_tracker.file_handlersimportload_human_codinghuman_results=load_human_coding(
"human_coding.csv",
doc_id_col="Media Title",
quote_col="Excerpt Copy",
range_col="Excerpt Range",
construct_col="Codes Applied Combined",
)

The defaults are designed for Dedoose-style excerpt exports. For other sources, pass the column names used by your file. The values in doc_id_col should match the document IDs produced by the LLM run.

Comparing Results

fromllm_tracker.comparisonimport (
LLMTrackerComparer,
compute_summary_tables,
format_concatenated,
format_weighted_summary,
)
comparer=LLMTrackerComparer(
api_key=api_key,
match_model="google/gemini-3-flash-preview",
)
comparison_table=comparer.compare_results(
human_results=human_results,
llm_results=results_llm,
output_dir="comparison_run",
)
per_doc, pooled, weighted=compute_summary_tables(comparison_table)
format_concatenated(pooled)
format_weighted_summary(weighted)

The comparison table contains one row per matched, human-only, or LLM-only construct instance. The human coding is treated as the reference set for classification metrics.

Quote Matching

Quote indices are recovered with exact matching by default. Fuzzy quote matching is available but off by default:

analyzer=LLMTrackerAnalyzer(
api_key=api_key,
fuzzy_quote_matching=True,
)

Use fuzzy matching when quotes may differ slightly from the source text due to spacing, punctuation, or small transcription differences.

Retry Failed Documents

Analyzer runs save error records for failed documents. You can retry them later:

recovered_results, recovered_metadata, remaining_errors=analyzer.retry_errors(
output_dir="LLM_coding_2026-05-20_120000",
codebook_path="codebook.json",
)

Tutorial

See tutorial.ipynb for a fuller walkthrough using the sample data and codebook. Open In Colab

Testing

Run the test suite with:

poetry run pytest

The tests avoid real API calls and focus on package behavior, file handling, comparison logic, configuration, and parsing.

Data Privacy

This package sends text to an LLM API during analysis and matching. Do not send identifiable or sensitive data unless it has been properly anonymized and your API provider's data handling policy is appropriate for your use case.

About

package for using LLMs in assessment of psychological constructs in text.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Feel free to use, but this is in beta version and is still being tested.

LLM Tracker

LLM Tracker is a Python package for identifying psychological constructs in text data (e.g., interviews, social media posts, chatbot interactions) and comparing LLM-coded results against human-coded results.

The package supports:

  • Access to thousands of remote or local LLMs through Mozilla's any-llm API (OpenRouter, Azure, Ollama, etc.)
  • Detect every instance of a construct (e.g., anxiety) and return the verbatim quote (e.g., "I'm worried about my cousin")
  • Comparing human and LLM codings at the quote level (using LLMs to match human and LLM quotes)
  • Computing inter-rater reliability metrics (Kappa, ICC, PABAK) and classification metrics (sensitivity, precision, F1, and PR AUC) and returning summary tables
  • Automatically retrying submissions when LLM outputs are not parseable
  • Saving analyzer outputs, metadata, and retryable error records
  • Flexible loading of csv, txt, docx and preprocessing of dedoose human coding to match LLM coding dataframes.
  • Many new features coming soon: visualizations, automated prompt engineering, and more!

Please cite this if you use this package:

Low, D., Mair, P., Nock, M., & Ghosh, S. (2025). Text Psychometrics: Assessing Psychological Constructs in Text Using Natural Language Processing. PsyArxiv. https://osf.io/preprints/psyarxiv/9rdux_v4

Installation

Install dependencies with Poetry:

poetry install

For tutorial extras such as corpus summaries:

poetry install --with tutorials

API Key

LLM Tracker uses any-llm API for LLM calls, which can run the most common APIs (OpenRouter, Azure, Ollama). For instance, you can add an OpenRouter API key by adding a few dollars here https://openrouter.ai/. Each LLM call tends to cost a a fraction of a cent (see cost for specific models on OpenRouter). Provide an API key directly:

api_key="your-openrouter-key"

or set it in the environment:

export OPENROUTER_API_KEY="your-openrouter-key"

You can also pass the path to a .env file containing:

OPENROUTER_API_KEY=your-openrouter-key

Basic LLM Coding

fromllm_trackerimportLLMTrackerAnalyzeranalyzer=LLMTrackerAnalyzer(
api_key=api_key,
model_name="google/gemini-3-flash-preview",
)
results_llm, metadata_llm, errors_llm=analyzer.analyze_csv(
csv_path="sample_data.csv",
codebook_path="codebook.json",
text_column="post",
subreddit_column="subreddit",
author_column="author",
output_dir="LLM_coding",
)

For a directory of supported document files:

results_llm, metadata_llm, errors_llm=analyzer.analyze_directory(
input_dir="documents",
codebook_path="codebook.json",
output_dir="LLM_coding",
)

Directory mode supports .txt and .csv files. Each file becomes one document.

Human Coding Input

Human coding is loaded into memory and passed directly to the comparer:

fromllm_tracker.file_handlersimportload_human_codinghuman_results=load_human_coding(
"human_coding.csv",
doc_id_col="Media Title",
quote_col="Excerpt Copy",
range_col="Excerpt Range",
construct_col="Codes Applied Combined",
)

The defaults are designed for Dedoose-style excerpt exports. For other sources, pass the column names used by your file. The values in doc_id_col should match the document IDs produced by the LLM run.

Comparing Results

fromllm_tracker.comparisonimport (
LLMTrackerComparer,
compute_summary_tables,
format_concatenated,
format_weighted_summary,
)
comparer=LLMTrackerComparer(
api_key=api_key,
match_model="google/gemini-3-flash-preview",
)
comparison_table=comparer.compare_results(
human_results=human_results,
llm_results=results_llm,
output_dir="comparison_run",
)
per_doc, pooled, weighted=compute_summary_tables(comparison_table)
format_concatenated(pooled)
format_weighted_summary(weighted)

The comparison table contains one row per matched, human-only, or LLM-only construct instance. The human coding is treated as the reference set for classification metrics.

Quote Matching

Quote indices are recovered with exact matching by default. Fuzzy quote matching is available but off by default:

analyzer=LLMTrackerAnalyzer(
api_key=api_key,
fuzzy_quote_matching=True,
)

Use fuzzy matching when quotes may differ slightly from the source text due to spacing, punctuation, or small transcription differences.

Retry Failed Documents

Analyzer runs save error records for failed documents. You can retry them later:

recovered_results, recovered_metadata, remaining_errors=analyzer.retry_errors(
output_dir="LLM_coding_2026-05-20_120000",
codebook_path="codebook.json",
)

Tutorial

See tutorial.ipynb for a fuller walkthrough using the sample data and codebook. Open In Colab

Testing

Run the test suite with:

poetry run pytest

The tests avoid real API calls and focus on package behavior, file handling, comparison logic, configuration, and parsing.

Data Privacy

This package sends text to an LLM API during analysis and matching. Do not send identifiable or sensitive data unless it has been properly anonymized and your API provider's data handling policy is appropriate for your use case.

About

package for using LLMs in assessment of psychological constructs in text.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Feel free to use, but this is in beta version and is still being tested.

LLM Tracker

LLM Tracker is a Python package for identifying psychological constructs in text data (e.g., interviews, social media posts, chatbot interactions) and comparing LLM-coded results against human-coded results.

The package supports:

  • Access to thousands of remote or local LLMs through Mozilla's any-llm API (OpenRouter, Azure, Ollama, etc.)
  • Detect every instance of a construct (e.g., anxiety) and return the verbatim quote (e.g., "I'm worried about my cousin")
  • Comparing human and LLM codings at the quote level (using LLMs to match human and LLM quotes)
  • Computing inter-rater reliability metrics (Kappa, ICC, PABAK) and classification metrics (sensitivity, precision, F1, and PR AUC) and returning summary tables
  • Automatically retrying submissions when LLM outputs are not parseable
  • Saving analyzer outputs, metadata, and retryable error records
  • Flexible loading of csv, txt, docx and preprocessing of dedoose human coding to match LLM coding dataframes.
  • Many new features coming soon: visualizations, automated prompt engineering, and more!

Please cite this if you use this package:

Low, D., Mair, P., Nock, M., & Ghosh, S. (2025). Text Psychometrics: Assessing Psychological Constructs in Text Using Natural Language Processing. PsyArxiv. https://osf.io/preprints/psyarxiv/9rdux_v4

Installation

Install dependencies with Poetry:

poetry install

For tutorial extras such as corpus summaries:

poetry install --with tutorials

API Key

LLM Tracker uses any-llm API for LLM calls, which can run the most common APIs (OpenRouter, Azure, Ollama). For instance, you can add an OpenRouter API key by adding a few dollars here https://openrouter.ai/. Each LLM call tends to cost a a fraction of a cent (see cost for specific models on OpenRouter). Provide an API key directly:

api_key="your-openrouter-key"

or set it in the environment:

export OPENROUTER_API_KEY="your-openrouter-key"

You can also pass the path to a .env file containing:

OPENROUTER_API_KEY=your-openrouter-key

Basic LLM Coding

fromllm_trackerimportLLMTrackerAnalyzeranalyzer=LLMTrackerAnalyzer(
api_key=api_key,
model_name="google/gemini-3-flash-preview",
)
results_llm, metadata_llm, errors_llm=analyzer.analyze_csv(
csv_path="sample_data.csv",
codebook_path="codebook.json",
text_column="post",
subreddit_column="subreddit",
author_column="author",
output_dir="LLM_coding",
)

For a directory of supported document files:

results_llm, metadata_llm, errors_llm=analyzer.analyze_directory(
input_dir="documents",
codebook_path="codebook.json",
output_dir="LLM_coding",
)

Directory mode supports .txt and .csv files. Each file becomes one document.

Human Coding Input

Human coding is loaded into memory and passed directly to the comparer:

fromllm_tracker.file_handlersimportload_human_codinghuman_results=load_human_coding(
"human_coding.csv",
doc_id_col="Media Title",
quote_col="Excerpt Copy",
range_col="Excerpt Range",
construct_col="Codes Applied Combined",
)

The defaults are designed for Dedoose-style excerpt exports. For other sources, pass the column names used by your file. The values in doc_id_col should match the document IDs produced by the LLM run.

Comparing Results

fromllm_tracker.comparisonimport (
LLMTrackerComparer,
compute_summary_tables,
format_concatenated,
format_weighted_summary,
)
comparer=LLMTrackerComparer(
api_key=api_key,
match_model="google/gemini-3-flash-preview",
)
comparison_table=comparer.compare_results(
human_results=human_results,
llm_results=results_llm,
output_dir="comparison_run",
)
per_doc, pooled, weighted=compute_summary_tables(comparison_table)
format_concatenated(pooled)
format_weighted_summary(weighted)

The comparison table contains one row per matched, human-only, or LLM-only construct instance. The human coding is treated as the reference set for classification metrics.

Quote Matching

Quote indices are recovered with exact matching by default. Fuzzy quote matching is available but off by default:

analyzer=LLMTrackerAnalyzer(
api_key=api_key,
fuzzy_quote_matching=True,
)

Use fuzzy matching when quotes may differ slightly from the source text due to spacing, punctuation, or small transcription differences.

Retry Failed Documents

Analyzer runs save error records for failed documents. You can retry them later:

recovered_results, recovered_metadata, remaining_errors=analyzer.retry_errors(
output_dir="LLM_coding_2026-05-20_120000",
codebook_path="codebook.json",
)

Tutorial

See tutorial.ipynb for a fuller walkthrough using the sample data and codebook. Open In Colab

Testing

Run the test suite with:

poetry run pytest

The tests avoid real API calls and focus on package behavior, file handling, comparison logic, configuration, and parsing.

Data Privacy

This package sends text to an LLM API during analysis and matching. Do not send identifiable or sensitive data unless it has been properly anonymized and your API provider's data handling policy is appropriate for your use case.

About

package for using LLMs in assessment of psychological constructs in text.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Feel free to use, but this is in beta version and is still being tested.

LLM Tracker

LLM Tracker is a Python package for identifying psychological constructs in text data (e.g., interviews, social media posts, chatbot interactions) and comparing LLM-coded results against human-coded results.

The package supports:

  • Access to thousands of remote or local LLMs through Mozilla's any-llm API (OpenRouter, Azure, Ollama, etc.)
  • Detect every instance of a construct (e.g., anxiety) and return the verbatim quote (e.g., "I'm worried about my cousin")
  • Comparing human and LLM codings at the quote level (using LLMs to match human and LLM quotes)
  • Computing inter-rater reliability metrics (Kappa, ICC, PABAK) and classification metrics (sensitivity, precision, F1, and PR AUC) and returning summary tables
  • Automatically retrying submissions when LLM outputs are not parseable
  • Saving analyzer outputs, metadata, and retryable error records
  • Flexible loading of csv, txt, docx and preprocessing of dedoose human coding to match LLM coding dataframes.
  • Many new features coming soon: visualizations, automated prompt engineering, and more!

Please cite this if you use this package:

Low, D., Mair, P., Nock, M., & Ghosh, S. (2025). Text Psychometrics: Assessing Psychological Constructs in Text Using Natural Language Processing. PsyArxiv. https://osf.io/preprints/psyarxiv/9rdux_v4

Installation

Install dependencies with Poetry:

poetry install

For tutorial extras such as corpus summaries:

poetry install --with tutorials

API Key

LLM Tracker uses any-llm API for LLM calls, which can run the most common APIs (OpenRouter, Azure, Ollama). For instance, you can add an OpenRouter API key by adding a few dollars here https://openrouter.ai/. Each LLM call tends to cost a a fraction of a cent (see cost for specific models on OpenRouter). Provide an API key directly:

api_key="your-openrouter-key"

or set it in the environment:

export OPENROUTER_API_KEY="your-openrouter-key"

You can also pass the path to a .env file containing:

OPENROUTER_API_KEY=your-openrouter-key

Basic LLM Coding

fromllm_trackerimportLLMTrackerAnalyzeranalyzer=LLMTrackerAnalyzer(
api_key=api_key,
model_name="google/gemini-3-flash-preview",
)
results_llm, metadata_llm, errors_llm=analyzer.analyze_csv(
csv_path="sample_data.csv",
codebook_path="codebook.json",
text_column="post",
subreddit_column="subreddit",
author_column="author",
output_dir="LLM_coding",
)

For a directory of supported document files:

results_llm, metadata_llm, errors_llm=analyzer.analyze_directory(
input_dir="documents",
codebook_path="codebook.json",
output_dir="LLM_coding",
)

Directory mode supports .txt and .csv files. Each file becomes one document.

Human Coding Input

Human coding is loaded into memory and passed directly to the comparer:

fromllm_tracker.file_handlersimportload_human_codinghuman_results=load_human_coding(
"human_coding.csv",
doc_id_col="Media Title",
quote_col="Excerpt Copy",
range_col="Excerpt Range",
construct_col="Codes Applied Combined",
)

The defaults are designed for Dedoose-style excerpt exports. For other sources, pass the column names used by your file. The values in doc_id_col should match the document IDs produced by the LLM run.

Comparing Results

fromllm_tracker.comparisonimport (
LLMTrackerComparer,
compute_summary_tables,
format_concatenated,
format_weighted_summary,
)
comparer=LLMTrackerComparer(
api_key=api_key,
match_model="google/gemini-3-flash-preview",
)
comparison_table=comparer.compare_results(
human_results=human_results,
llm_results=results_llm,
output_dir="comparison_run",
)
per_doc, pooled, weighted=compute_summary_tables(comparison_table)
format_concatenated(pooled)
format_weighted_summary(weighted)

The comparison table contains one row per matched, human-only, or LLM-only construct instance. The human coding is treated as the reference set for classification metrics.

Quote Matching

Quote indices are recovered with exact matching by default. Fuzzy quote matching is available but off by default:

analyzer=LLMTrackerAnalyzer(
api_key=api_key,
fuzzy_quote_matching=True,
)

Use fuzzy matching when quotes may differ slightly from the source text due to spacing, punctuation, or small transcription differences.

Retry Failed Documents

Analyzer runs save error records for failed documents. You can retry them later:

recovered_results, recovered_metadata, remaining_errors=analyzer.retry_errors(
output_dir="LLM_coding_2026-05-20_120000",
codebook_path="codebook.json",
)

Tutorial

See tutorial.ipynb for a fuller walkthrough using the sample data and codebook. Open In Colab

Testing

Run the test suite with:

poetry run pytest

The tests avoid real API calls and focus on package behavior, file handling, comparison logic, configuration, and parsing.

Data Privacy

This package sends text to an LLM API during analysis and matching. Do not send identifiable or sensitive data unless it has been properly anonymized and your API provider's data handling policy is appropriate for your use case.

About

package for using LLMs in assessment of psychological constructs in text.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Feel free to use, but this is in beta version and is still being tested.

LLM Tracker

LLM Tracker is a Python package for identifying psychological constructs in text data (e.g., interviews, social media posts, chatbot interactions) and comparing LLM-coded results against human-coded results.

The package supports:

  • Access to thousands of remote or local LLMs through Mozilla's any-llm API (OpenRouter, Azure, Ollama, etc.)
  • Detect every instance of a construct (e.g., anxiety) and return the verbatim quote (e.g., "I'm worried about my cousin")
  • Comparing human and LLM codings at the quote level (using LLMs to match human and LLM quotes)
  • Computing inter-rater reliability metrics (Kappa, ICC, PABAK) and classification metrics (sensitivity, precision, F1, and PR AUC) and returning summary tables
  • Automatically retrying submissions when LLM outputs are not parseable
  • Saving analyzer outputs, metadata, and retryable error records
  • Flexible loading of csv, txt, docx and preprocessing of dedoose human coding to match LLM coding dataframes.
  • Many new features coming soon: visualizations, automated prompt engineering, and more!

Please cite this if you use this package:

Low, D., Mair, P., Nock, M., & Ghosh, S. (2025). Text Psychometrics: Assessing Psychological Constructs in Text Using Natural Language Processing. PsyArxiv. https://osf.io/preprints/psyarxiv/9rdux_v4

Installation

Install dependencies with Poetry:

poetry install

For tutorial extras such as corpus summaries:

poetry install --with tutorials

API Key

LLM Tracker uses any-llm API for LLM calls, which can run the most common APIs (OpenRouter, Azure, Ollama). For instance, you can add an OpenRouter API key by adding a few dollars here https://openrouter.ai/. Each LLM call tends to cost a a fraction of a cent (see cost for specific models on OpenRouter). Provide an API key directly:

api_key="your-openrouter-key"

or set it in the environment:

export OPENROUTER_API_KEY="your-openrouter-key"

You can also pass the path to a .env file containing:

OPENROUTER_API_KEY=your-openrouter-key

Basic LLM Coding

fromllm_trackerimportLLMTrackerAnalyzeranalyzer=LLMTrackerAnalyzer(
api_key=api_key,
model_name="google/gemini-3-flash-preview",
)
results_llm, metadata_llm, errors_llm=analyzer.analyze_csv(
csv_path="sample_data.csv",
codebook_path="codebook.json",
text_column="post",
subreddit_column="subreddit",
author_column="author",
output_dir="LLM_coding",
)

For a directory of supported document files:

results_llm, metadata_llm, errors_llm=analyzer.analyze_directory(
input_dir="documents",
codebook_path="codebook.json",
output_dir="LLM_coding",
)

Directory mode supports .txt and .csv files. Each file becomes one document.

Human Coding Input

Human coding is loaded into memory and passed directly to the comparer:

fromllm_tracker.file_handlersimportload_human_codinghuman_results=load_human_coding(
"human_coding.csv",
doc_id_col="Media Title",
quote_col="Excerpt Copy",
range_col="Excerpt Range",
construct_col="Codes Applied Combined",
)

The defaults are designed for Dedoose-style excerpt exports. For other sources, pass the column names used by your file. The values in doc_id_col should match the document IDs produced by the LLM run.

Comparing Results

fromllm_tracker.comparisonimport (
LLMTrackerComparer,
compute_summary_tables,
format_concatenated,
format_weighted_summary,
)
comparer=LLMTrackerComparer(
api_key=api_key,
match_model="google/gemini-3-flash-preview",
)
comparison_table=comparer.compare_results(
human_results=human_results,
llm_results=results_llm,
output_dir="comparison_run",
)
per_doc, pooled, weighted=compute_summary_tables(comparison_table)
format_concatenated(pooled)
format_weighted_summary(weighted)

The comparison table contains one row per matched, human-only, or LLM-only construct instance. The human coding is treated as the reference set for classification metrics.

Quote Matching

Quote indices are recovered with exact matching by default. Fuzzy quote matching is available but off by default:

analyzer=LLMTrackerAnalyzer(
api_key=api_key,
fuzzy_quote_matching=True,
)

Use fuzzy matching when quotes may differ slightly from the source text due to spacing, punctuation, or small transcription differences.

Retry Failed Documents

Analyzer runs save error records for failed documents. You can retry them later:

recovered_results, recovered_metadata, remaining_errors=analyzer.retry_errors(
output_dir="LLM_coding_2026-05-20_120000",
codebook_path="codebook.json",
)

Tutorial

See tutorial.ipynb for a fuller walkthrough using the sample data and codebook. Open In Colab

Testing

Run the test suite with:

poetry run pytest

The tests avoid real API calls and focus on package behavior, file handling, comparison logic, configuration, and parsing.

Data Privacy

This package sends text to an LLM API during analysis and matching. Do not send identifiable or sensitive data unless it has been properly anonymized and your API provider's data handling policy is appropriate for your use case.

About

package for using LLMs in assessment of psychological constructs in text.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Feel free to use, but this is in beta version and is still being tested.

LLM Tracker

LLM Tracker is a Python package for identifying psychological constructs in text data (e.g., interviews, social media posts, chatbot interactions) and comparing LLM-coded results against human-coded results.

The package supports:

  • Access to thousands of remote or local LLMs through Mozilla's any-llm API (OpenRouter, Azure, Ollama, etc.)
  • Detect every instance of a construct (e.g., anxiety) and return the verbatim quote (e.g., "I'm worried about my cousin")
  • Comparing human and LLM codings at the quote level (using LLMs to match human and LLM quotes)
  • Computing inter-rater reliability metrics (Kappa, ICC, PABAK) and classification metrics (sensitivity, precision, F1, and PR AUC) and returning summary tables
  • Automatically retrying submissions when LLM outputs are not parseable
  • Saving analyzer outputs, metadata, and retryable error records
  • Flexible loading of csv, txt, docx and preprocessing of dedoose human coding to match LLM coding dataframes.
  • Many new features coming soon: visualizations, automated prompt engineering, and more!

Please cite this if you use this package:

Low, D., Mair, P., Nock, M., & Ghosh, S. (2025). Text Psychometrics: Assessing Psychological Constructs in Text Using Natural Language Processing. PsyArxiv. https://osf.io/preprints/psyarxiv/9rdux_v4

Installation

Install dependencies with Poetry:

poetry install

For tutorial extras such as corpus summaries:

poetry install --with tutorials

API Key

LLM Tracker uses any-llm API for LLM calls, which can run the most common APIs (OpenRouter, Azure, Ollama). For instance, you can add an OpenRouter API key by adding a few dollars here https://openrouter.ai/. Each LLM call tends to cost a a fraction of a cent (see cost for specific models on OpenRouter). Provide an API key directly:

api_key="your-openrouter-key"

or set it in the environment:

export OPENROUTER_API_KEY="your-openrouter-key"

You can also pass the path to a .env file containing:

OPENROUTER_API_KEY=your-openrouter-key

Basic LLM Coding

fromllm_trackerimportLLMTrackerAnalyzeranalyzer=LLMTrackerAnalyzer(
api_key=api_key,
model_name="google/gemini-3-flash-preview",
)
results_llm, metadata_llm, errors_llm=analyzer.analyze_csv(
csv_path="sample_data.csv",
codebook_path="codebook.json",
text_column="post",
subreddit_column="subreddit",
author_column="author",
output_dir="LLM_coding",
)

For a directory of supported document files:

results_llm, metadata_llm, errors_llm=analyzer.analyze_directory(
input_dir="documents",
codebook_path="codebook.json",
output_dir="LLM_coding",
)

Directory mode supports .txt and .csv files. Each file becomes one document.

Human Coding Input

Human coding is loaded into memory and passed directly to the comparer:

fromllm_tracker.file_handlersimportload_human_codinghuman_results=load_human_coding(
"human_coding.csv",
doc_id_col="Media Title",
quote_col="Excerpt Copy",
range_col="Excerpt Range",
construct_col="Codes Applied Combined",
)

The defaults are designed for Dedoose-style excerpt exports. For other sources, pass the column names used by your file. The values in doc_id_col should match the document IDs produced by the LLM run.

Comparing Results

fromllm_tracker.comparisonimport (
LLMTrackerComparer,
compute_summary_tables,
format_concatenated,
format_weighted_summary,
)
comparer=LLMTrackerComparer(
api_key=api_key,
match_model="google/gemini-3-flash-preview",
)
comparison_table=comparer.compare_results(
human_results=human_results,
llm_results=results_llm,
output_dir="comparison_run",
)
per_doc, pooled, weighted=compute_summary_tables(comparison_table)
format_concatenated(pooled)
format_weighted_summary(weighted)

The comparison table contains one row per matched, human-only, or LLM-only construct instance. The human coding is treated as the reference set for classification metrics.

Quote Matching

Quote indices are recovered with exact matching by default. Fuzzy quote matching is available but off by default:

analyzer=LLMTrackerAnalyzer(
api_key=api_key,
fuzzy_quote_matching=True,
)

Use fuzzy matching when quotes may differ slightly from the source text due to spacing, punctuation, or small transcription differences.

Retry Failed Documents

Analyzer runs save error records for failed documents. You can retry them later:

recovered_results, recovered_metadata, remaining_errors=analyzer.retry_errors(
output_dir="LLM_coding_2026-05-20_120000",
codebook_path="codebook.json",
)

Tutorial

See tutorial.ipynb for a fuller walkthrough using the sample data and codebook. Open In Colab

Testing

Run the test suite with:

poetry run pytest

The tests avoid real API calls and focus on package behavior, file handling, comparison logic, configuration, and parsing.

Data Privacy

This package sends text to an LLM API during analysis and matching. Do not send identifiable or sensitive data unless it has been properly anonymized and your API provider's data handling policy is appropriate for your use case.

About

package for using LLMs in assessment of psychological constructs in text.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Feel free to use, but this is in beta version and is still being tested.

LLM Tracker

LLM Tracker is a Python package for identifying psychological constructs in text data (e.g., interviews, social media posts, chatbot interactions) and comparing LLM-coded results against human-coded results.

The package supports:

  • Access to thousands of remote or local LLMs through Mozilla's any-llm API (OpenRouter, Azure, Ollama, etc.)
  • Detect every instance of a construct (e.g., anxiety) and return the verbatim quote (e.g., "I'm worried about my cousin")
  • Comparing human and LLM codings at the quote level (using LLMs to match human and LLM quotes)
  • Computing inter-rater reliability metrics (Kappa, ICC, PABAK) and classification metrics (sensitivity, precision, F1, and PR AUC) and returning summary tables
  • Automatically retrying submissions when LLM outputs are not parseable
  • Saving analyzer outputs, metadata, and retryable error records
  • Flexible loading of csv, txt, docx and preprocessing of dedoose human coding to match LLM coding dataframes.
  • Many new features coming soon: visualizations, automated prompt engineering, and more!

Please cite this if you use this package:

Low, D., Mair, P., Nock, M., & Ghosh, S. (2025). Text Psychometrics: Assessing Psychological Constructs in Text Using Natural Language Processing. PsyArxiv. https://osf.io/preprints/psyarxiv/9rdux_v4

Installation

Install dependencies with Poetry:

poetry install

For tutorial extras such as corpus summaries:

poetry install --with tutorials

API Key

LLM Tracker uses any-llm API for LLM calls, which can run the most common APIs (OpenRouter, Azure, Ollama). For instance, you can add an OpenRouter API key by adding a few dollars here https://openrouter.ai/. Each LLM call tends to cost a a fraction of a cent (see cost for specific models on OpenRouter). Provide an API key directly:

api_key="your-openrouter-key"

or set it in the environment:

export OPENROUTER_API_KEY="your-openrouter-key"

You can also pass the path to a .env file containing:

OPENROUTER_API_KEY=your-openrouter-key

Basic LLM Coding

fromllm_trackerimportLLMTrackerAnalyzeranalyzer=LLMTrackerAnalyzer(
api_key=api_key,
model_name="google/gemini-3-flash-preview",
)
results_llm, metadata_llm, errors_llm=analyzer.analyze_csv(
csv_path="sample_data.csv",
codebook_path="codebook.json",
text_column="post",
subreddit_column="subreddit",
author_column="author",
output_dir="LLM_coding",
)

For a directory of supported document files:

results_llm, metadata_llm, errors_llm=analyzer.analyze_directory(
input_dir="documents",
codebook_path="codebook.json",
output_dir="LLM_coding",
)

Directory mode supports .txt and .csv files. Each file becomes one document.

Human Coding Input

Human coding is loaded into memory and passed directly to the comparer:

fromllm_tracker.file_handlersimportload_human_codinghuman_results=load_human_coding(
"human_coding.csv",
doc_id_col="Media Title",
quote_col="Excerpt Copy",
range_col="Excerpt Range",
construct_col="Codes Applied Combined",
)

The defaults are designed for Dedoose-style excerpt exports. For other sources, pass the column names used by your file. The values in doc_id_col should match the document IDs produced by the LLM run.

Comparing Results

fromllm_tracker.comparisonimport (
LLMTrackerComparer,
compute_summary_tables,
format_concatenated,
format_weighted_summary,
)
comparer=LLMTrackerComparer(
api_key=api_key,
match_model="google/gemini-3-flash-preview",
)
comparison_table=comparer.compare_results(
human_results=human_results,
llm_results=results_llm,
output_dir="comparison_run",
)
per_doc, pooled, weighted=compute_summary_tables(comparison_table)
format_concatenated(pooled)
format_weighted_summary(weighted)

The comparison table contains one row per matched, human-only, or LLM-only construct instance. The human coding is treated as the reference set for classification metrics.

Quote Matching

Quote indices are recovered with exact matching by default. Fuzzy quote matching is available but off by default:

analyzer=LLMTrackerAnalyzer(
api_key=api_key,
fuzzy_quote_matching=True,
)

Use fuzzy matching when quotes may differ slightly from the source text due to spacing, punctuation, or small transcription differences.

Retry Failed Documents

Analyzer runs save error records for failed documents. You can retry them later:

recovered_results, recovered_metadata, remaining_errors=analyzer.retry_errors(
output_dir="LLM_coding_2026-05-20_120000",
codebook_path="codebook.json",
)

Tutorial

See tutorial.ipynb for a fuller walkthrough using the sample data and codebook. Open In Colab

Testing

Run the test suite with:

poetry run pytest

The tests avoid real API calls and focus on package behavior, file handling, comparison logic, configuration, and parsing.

Data Privacy

This package sends text to an LLM API during analysis and matching. Do not send identifiable or sensitive data unless it has been properly anonymized and your API provider's data handling policy is appropriate for your use case.

About

package for using LLMs in assessment of psychological constructs in text.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Feel free to use, but this is in beta version and is still being tested.

LLM Tracker

LLM Tracker is a Python package for identifying psychological constructs in text data (e.g., interviews, social media posts, chatbot interactions) and comparing LLM-coded results against human-coded results.

The package supports:

  • Access to thousands of remote or local LLMs through Mozilla's any-llm API (OpenRouter, Azure, Ollama, etc.)
  • Detect every instance of a construct (e.g., anxiety) and return the verbatim quote (e.g., "I'm worried about my cousin")
  • Comparing human and LLM codings at the quote level (using LLMs to match human and LLM quotes)
  • Computing inter-rater reliability metrics (Kappa, ICC, PABAK) and classification metrics (sensitivity, precision, F1, and PR AUC) and returning summary tables
  • Automatically retrying submissions when LLM outputs are not parseable
  • Saving analyzer outputs, metadata, and retryable error records
  • Flexible loading of csv, txt, docx and preprocessing of dedoose human coding to match LLM coding dataframes.
  • Many new features coming soon: visualizations, automated prompt engineering, and more!

Please cite this if you use this package:

Low, D., Mair, P., Nock, M., & Ghosh, S. (2025). Text Psychometrics: Assessing Psychological Constructs in Text Using Natural Language Processing. PsyArxiv. https://osf.io/preprints/psyarxiv/9rdux_v4

Installation

Install dependencies with Poetry:

poetry install

For tutorial extras such as corpus summaries:

poetry install --with tutorials

API Key

LLM Tracker uses any-llm API for LLM calls, which can run the most common APIs (OpenRouter, Azure, Ollama). For instance, you can add an OpenRouter API key by adding a few dollars here https://openrouter.ai/. Each LLM call tends to cost a a fraction of a cent (see cost for specific models on OpenRouter). Provide an API key directly:

api_key="your-openrouter-key"

or set it in the environment:

export OPENROUTER_API_KEY="your-openrouter-key"

You can also pass the path to a .env file containing:

OPENROUTER_API_KEY=your-openrouter-key

Basic LLM Coding

fromllm_trackerimportLLMTrackerAnalyzeranalyzer=LLMTrackerAnalyzer(
api_key=api_key,
model_name="google/gemini-3-flash-preview",
)
results_llm, metadata_llm, errors_llm=analyzer.analyze_csv(
csv_path="sample_data.csv",
codebook_path="codebook.json",
text_column="post",
subreddit_column="subreddit",
author_column="author",
output_dir="LLM_coding",
)

For a directory of supported document files:

results_llm, metadata_llm, errors_llm=analyzer.analyze_directory(
input_dir="documents",
codebook_path="codebook.json",
output_dir="LLM_coding",
)

Directory mode supports .txt and .csv files. Each file becomes one document.

Human Coding Input

Human coding is loaded into memory and passed directly to the comparer:

fromllm_tracker.file_handlersimportload_human_codinghuman_results=load_human_coding(
"human_coding.csv",
doc_id_col="Media Title",
quote_col="Excerpt Copy",
range_col="Excerpt Range",
construct_col="Codes Applied Combined",
)

The defaults are designed for Dedoose-style excerpt exports. For other sources, pass the column names used by your file. The values in doc_id_col should match the document IDs produced by the LLM run.

Comparing Results

fromllm_tracker.comparisonimport (
LLMTrackerComparer,
compute_summary_tables,
format_concatenated,
format_weighted_summary,
)
comparer=LLMTrackerComparer(
api_key=api_key,
match_model="google/gemini-3-flash-preview",
)
comparison_table=comparer.compare_results(
human_results=human_results,
llm_results=results_llm,
output_dir="comparison_run",
)
per_doc, pooled, weighted=compute_summary_tables(comparison_table)
format_concatenated(pooled)
format_weighted_summary(weighted)

The comparison table contains one row per matched, human-only, or LLM-only construct instance. The human coding is treated as the reference set for classification metrics.

Quote Matching

Quote indices are recovered with exact matching by default. Fuzzy quote matching is available but off by default:

analyzer=LLMTrackerAnalyzer(
api_key=api_key,
fuzzy_quote_matching=True,
)

Use fuzzy matching when quotes may differ slightly from the source text due to spacing, punctuation, or small transcription differences.

Retry Failed Documents

Analyzer runs save error records for failed documents. You can retry them later:

recovered_results, recovered_metadata, remaining_errors=analyzer.retry_errors(
output_dir="LLM_coding_2026-05-20_120000",
codebook_path="codebook.json",
)

Tutorial

See tutorial.ipynb for a fuller walkthrough using the sample data and codebook. Open In Colab

Testing

Run the test suite with:

poetry run pytest

The tests avoid real API calls and focus on package behavior, file handling, comparison logic, configuration, and parsing.

Data Privacy

This package sends text to an LLM API during analysis and matching. Do not send identifiable or sensitive data unless it has been properly anonymized and your API provider's data handling policy is appropriate for your use case.

About

package for using LLMs in assessment of psychological constructs in text.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages