This script automates the process of creating Anki flashcards from words saved in a Todoist project (or a CSV/text file). It fetches items, gets word definitions and example sentences using an LLM, and creates the notes directly in Anki via AnkiConnect.
- Todoist Integration: Pulls tasks from a specified Todoist project.
- Flexible Word Parsing: Extracts words from task titles like
{word},English {word}, or justword. - AI-Powered Definitions: Uses an LLM to get context-aware definitions for each word.
- AI-Powered Sentence Generation: Generates an example sentence for each word.
- Multi-language support: Build decks for any language (Estonian, Russian, …) via
--learning-language; control the language of definitions separately via--instruction-language. - Anki Card Generation: Creates cloze-deletion flashcards via AnkiConnect.
- Per-language subdecks: Every language gets its own subdeck —
sentence-mining::<language>(e.g.sentence-mining::english,sentence-mining::estonian). The root deck is an empty container. - Secure: Uses a
.envfile to keep your API keys safe. - Automation-Ready: Can be easily set up with a cron job to run automatically.
Follow these steps to set up and run the project.
git clone <repository-url>
cd <repository-name>It's highly recommended to use a virtual environment to manage project dependencies.
# Create a virtual environment named 'venv'
python3 -m venv venv
# Activate the virtual environment
# On macOS and Linux:
source venv/bin/activate
# On Windows:
# venv\Scripts\activateInstall the required Python libraries using pip.
pip install -r requirements.txtThe script loads your API keys from a .env file.
- Create a
.envfile in the project root by copying the example file:cp .env.example .env
- Open the
.envfile and add your API keys:The Nebius model is chosen at runtime with the requiredTODOIST_API_KEY="YOUR_TODOIST_API_KEY" NEBIUS_API_KEY="YOUR_NEBIUS_API_KEY"--modelflag (see below) — there is no env var or code default.
Make sure Anki is open and running, then execute main.py:
python main.py --model <model-id> \
[--source <todoist|csv|text_file>] \
[--csv-file <path>] \
[--text-file <path>] \
[--tags <tag1,tag2,...>] \
[--learning-language <code|name>] \
[--instruction-language <code|name>]Flags:
| Flag | Default | Description |
|---|---|---|
--model |
— | Required. Nebius model ID (e.g. openai/gpt-oss-120b) |
--source |
todoist |
Input source: todoist, csv, or text_file |
--csv-file |
— | Path to CSV file (required when --source csv) |
--text-file |
— | Path to text file (required when --source text_file) |
--tags / -t |
— | Comma-separated tags added to every note in the run |
--learning-language |
english |
Language of the word and generated example sentence. Accepts ISO 639-1 codes (en, et, ru) or full names. |
--instruction-language |
english |
Language of the definition/explanation. Accepts the same codes/names. |
Examples:
# Default: Todoist source, English
python main.py --model openai/gpt-oss-120b
# Estonian CSV → Estonian deck + English definitions (default)
python main.py --model openai/gpt-oss-120b \
--source csv --csv-file estonian_words.csv --learning-language et
# Estonian CSV with tags (e.g. a specific book)
python main.py --model openai/gpt-oss-120b \
--source csv --csv-file csv-files/my-estonian-book.csv \
--learning-language et \
--tags "Source::MyBook,Topic::Fiction,Type::Book"
# Estonian CSV → Estonian deck + Russian definitions
python main.py --model openai/gpt-oss-120b \
--source csv --csv-file estonian_words.csv \
--learning-language et --instruction-language ru
# English book import with tags
python main.py --model openai/gpt-oss-120b \
--source csv --csv-file my_book.csv \
--tags "Source::MyBook,Topic::History,Type::Book"Cards are created via AnkiConnect into sentence-mining::<language> subdecks (e.g. sentence-mining::english, sentence-mining::estonian). The root sentence-mining deck is an empty container.
CsvSentenceSource skips the header row and infers columns purely from count, so header text is a label only — get the column count right and any header works:
| Columns | Layout | Notes |
|---|---|---|
| 2 | word,context |
Recommended. Tags come from --tags on the CLI, applied to every row. This is what every file in csv-files/ uses. |
| 3 | id,word,context |
Adds a per-row id (logging only; CSV rows are never marked "complete"). |
| 4 | id,word,context,tags |
Adds a per-row tags column, comma-separated. Must be quoted if it contains more than one tag ("Tag1,Tag2") — an unquoted multi-tag cell overflows into extra columns and silently drops everything past the 4th. |
word can be a single word or a multi-word phrase/idiom (e.g. beat the bushes) — no special markup needed, it's used verbatim. Quote any context cell that contains a comma. Example (2-column form):
word,context
gruesome,"If you use it without paying careful attention, the result can be gruesome."
horse sense,"But also the horse sense to work things out on the fly."You can automate the script to run at regular intervals using a cron job.
-
Open your crontab file for editing:
crontab -e
-
Add a new line to schedule the job. The following example runs the script every day at 7:00 AM.
0 7 * * * /path/to/your/project/venv/bin/python /path/to/your/project/main.pyImportant:
- Replace
/path/to/your/project/with the absolute path to this project's directory. - The command specifies the Python executable inside the virtual environment (
venv/bin/python) to ensure the correct dependencies are used.
- Replace
-
Save and exit the crontab editor. The cron job is now active.
This script follows a layered architecture to separate concerns and make the code easier to maintain and extend (see AGENTS.md for the full breakdown).
- Domain layer (
domain/): core entities and interfaces —SourceSentence,SentenceSource,TaskCompletionHandler— independent of any specific data source. - Repositories layer (
repositories/): thin wrappers around external APIs —TodoistRepository,LLMRepository(Nebius),AnkiRepository(AnkiConnect) — with retry logic viatenacity. - Data sources layer (
datasources/): implementations ofSentenceSourceper input type —TodoistSentenceSource,CsvSentenceSource,TextFileSentenceSource. - Service layer:
llm_service.py,anki_service.py,word_processor.py— business logic, injected with repositories rather than constructing them. - Composition root:
main.pyparses CLI args and wires the above together;config.pyholds static config and secrets loaded from.env. - Anki integration:
anki_service.pycreates notes directly via AnkiConnect (no local.apkgpackaging), with cloze-deletion and definition→word card templates, plus duplicate-note handling described below.
The application implements a flexible tagging system for Anki notes, combining tags from multiple sources. This system utilizes a nested tag hierarchy (using ::) for better organization and leverages Anki's powerful filtering capabilities.
Recommended Tag Structure:
- Time:
Year::YYYY(e.g.,Year::2026),Month::MM(e.g.,Month::01), andInstructionLanguage::<Language>(e.g.,InstructionLanguage::English). These are all automatically generated. - Source Type:
Type::Book,Type::News,Type::Podcast, etc. (e.g.,Type::Bookfrom a CSV or text file,Type::Todoistfor Todoist tasks). - Specific Source:
Source::BookName,Source::NewspaperName,Source::PodcastName(e.g.,Source::Harry_Potter,Source::New_Yorker,Source::NPR_Podcast). This can be added via command-line arguments for batch processing. - Subject/Domain:
Topic::Tech,Topic::Finance,Topic::Literature,Topic::History, etc. - Functional Tags (User-Defined): These tags describe how the card behaves or its status.
Check: For cards that might have a typo, an incorrect definition, or require manual review.IdiomorPhrasalVerb: To categorize multi-word expressions.Critical: For words or phrases that are essential to know (e.g., for work, an exam).
How Tagging Works:
- Combination: Tags are collected from script-generated defaults, data source metadata (e.g., Todoist task labels, CSV
tagscolumn), and command-line arguments (--tagsor-t). - Deduplication: All collected tags are combined, and duplicates are automatically removed.
- Hierarchical Format: Anki's hierarchical tag format (e.g.,
Parent::Child) is used for better organization. - Benefits: Using a robust tagging system in Anki allows for flexible study. You can create "Filtered Decks" based on specific tags (e.g., to study only words from a particular book before a test) while keeping all your cards in one main deck for daily, efficient review.
Example Usage (Command Line):
python main.py --model openai/gpt-oss-120b --source csv --csv-file my_book.csv --tags "Source::MyBook,Topic::History,Type::Book"
python main.py --model openai/gpt-oss-120b --source text_file --text-file my_sentences.txt --tags "Source::Article_Title,Topic::Science,Check"
python main.py --model openai/gpt-oss-120b --source csv --csv-file csv-files/book-project-hail-mari-until-page62.csv --tags "Source::Project_Hail_Mary,Topic::SciFi,Type::Book"First, you need to give the bash script execute permissions:
chmod +x /path/to/your/project/sentence_miner_todoist.shReplace /path/to/your/project/ with the actual path to your project directory.
Before setting up the cron job, test that the script works correctly:
/path/to/your/project/sentence_miner_todoist.shThis should:
- Activate your virtual environment
- Run the main script with the todoist source
- Generate the Anki deck
- Log completion to
cron.log
Open your crontab file for editing:
crontab -eIf this is your first time, you may be asked to choose an editor. Select your preferred editor (nano is easiest for beginners).
Add one of the following lines to schedule your job:
0 7 * * * /path/to/your/project/sentence_miner_todoist.sh >> /path/to/your/project/cron.log 2>&1
0 */6 * * * /path/to/your/project/sentence_miner_todoist.sh >> /path/to/your/project/cron.log 2>&1
0 22 * * * /path/to/your/project/sentence_miner_todoist.sh >> /path/to/your/project/cron.log 2>&1
0 8,20 * * * /path/to/your/project/sentence_miner_todoist.sh >> /path/to/your/project/cron.log 2>&1
Important: Replace /path/to/your/project/ with your actual project path in both places.
- If using nano: Press
Ctrl+X, thenY, thenEnter - If using vim: Press
Esc, type:wq, thenEnter
Check that your cron job was added successfully:
crontab -lThis should display all your scheduled cron jobs, including the one you just added.
The cron time format is: minute hour day month day_of_week
* * * * * command
│ │ │ │ │
│ │ │ │ └─── Day of week (0-7, both 0 and 7 are Sunday)
│ │ │ └──────── Month (1-12)
│ │ └───────────── Day of month (1-31)
│ └────────────────── Hour (0-23)
└─────────────────────── Minute (0-59)
Examples:
0 7 * * *- Every day at 7:00 AM*/30 * * * *- Every 30 minutes0 */4 * * *- Every 4 hours0 9 * * 1- Every Monday at 9:00 AM0 0 1 * *- First day of every month at midnight
sudo systemctl status cronOn most systems:
grep CRON /var/log/syslogOr check your project's log file:
cat /path/to/your/project/cron.log-
Script doesn't run:
- Verify the script has execute permissions (
chmod +x) - Check that paths are absolute (not relative)
- Ensure the
.envfile exists in the project directory
- Verify the script has execute permissions (
-
Environment variables not loading:
- The script automatically changes to the project directory before running
- Make sure your
.envfile is in the project root
-
Virtual environment issues:
- Verify the venv exists at
venv/bin/activate - Test the script manually first
- Verify the venv exists at
To test if your cron job will run soon, you can temporarily set it to run in a few minutes:
# Get current time
date
# Edit crontab
crontab -e
# Add a test job that runs 2 minutes from now
# For example, if it's 14:30, add:
32 14 * * * /path/to/your/project/sentence_miner_todoist.sh >> /path/to/your/project/cron.log 2>&1
# Wait and check the log file
tail -f /path/to/your/project/cron.logIf you need to temporarily disable the cron job:
crontab -eThen add a # at the beginning of the line to comment it out:
# 0 7 * * * /path/to/your/project/sentence_miner_todoist.sh >> /path/to/your/project/cron.log 2>&1
To completely remove all cron jobs:
crontab -r(Use with caution! This removes ALL your cron jobs.)
-
Email Notifications: By default, cron sends email on errors. To disable:
MAILTO="" 0 7 * * * /path/to/your/project/sentence_miner_todoist.sh >> /path/to/your/project/cron.log 2>&1 -
Multiple Environments: If you have multiple Python projects, each should have its own bash script pointing to its own venv.
-
Backup Your Collection: Notes are written straight into Anki via AnkiConnect (no intermediate export file), so back up your Anki collection itself (e.g. via Anki's built-in backups or exporting a deck package from within Anki) rather than looking for generated files in this project.
