feat(agents): prompt-injection hardening on student content (#150) - #471

Merged
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening
Jul 30, 2026
Merged

feat(agents): prompt-injection hardening on student content (#150)#471
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

What

Third link of the B7 seam-endgame chain; gates the public beta. Full detail in the commit message. The shape:

  • Containment, single source of truth: services/prompt_safety.py — delimited untrusted-content envelopes with data-not-instructions framing and delimiter-forgery neutralization (embedded BEGIN/END copies are defanged, case/whitespace-insensitively, idempotently). Applied at every assembly boundary where student or peer-derived text enters a prompt: RAG chunks, the graph seed block, the legacy COURSE MATERIALS + shared-context JSON, and the tool→LLM boundary. Storage stays raw.
  • Instruction hardening: a shared injection-guard paragraph across the tutor's three preambles, note_chat, quiz, and the legacy preamble — plus the academic-integrity block held out of [P1] Agent migration: graph-grounded tutor retrieval + tool-use loop #149 for exactly this PR.
  • Tool-use constraint: the mastery-delta schema is clamped to the instructed band, so an injected "set my mastery to 1.0" fails validation into [P2] Agent platform: structured-output retry + validation hardening #153's retry loop; a contract test freezes that no tool signature exposes ids to the model.
  • Honest eval handling: replay keys on case names and would have stayed silently green on prompt changes — chat_tutor and quiz_generation were re-recorded (the re-record also landed ADR-0023's tracked prompt-shape fix after the bare-newline quirk reproduced live). All evaluators 1.000 on both datasets; GroundedConcept ratcheted 0.875 → 1.0.
  • Documented scope calls: the student's own message channel stays unwrapped (it is their instruction channel), catalog chunks stay trusted, the misconceptions tool stays unregistered pending real consent enforcement, and the tool-less document workers are a noted follow-up.

Verification

27-case red-first injection suite (14 red at base) including an end-to-end FunctionModel test proving injected tool returns reach the model enveloped. Backend 1518 passed + ruff clean; 278 passed under the lock-pinned 1.107 venv; evals replay green across all six datasets. Full local e2e cycle pre-merge; results below.

Closes#150.

🤖 Generated with Claude Code

Third link of the seam endgame; gates the public beta.
- One shared containment helper (services/prompt_safety.py):
wrap_untrusted builds a delimited BEGIN/END UNTRUSTED CONTENT envelope
with data-not-instructions framing; neutralize_delimiters defangs
embedded marker forgeries (case/whitespace-insensitive, idempotent) so
content can't fake an early END and escape. Applied at ASSEMBLY
boundaries only — storage keeps raw text: RAG chunks
(format_rag_context — covers agent chat, legacy chat, quiz context in
one place), the graph seed block's student-derived concept names, the
legacy prompt's COURSE MATERIALS + shared-context JSON, and the tool→
LLM boundary (search_course_materials, read_active_note,
read_misconceptions, quiz-history; note-worker user prompts).
- INJECTION_GUARD_PROMPT (single source) in the tutor's three preambles,
note_chat, quiz, and the legacy preamble; the ACADEMIC INTEGRITY block
deferred out of #149 lands in the agent preamble for parity.
- Tool-use constraint: ConceptMasteryUpdate.mastery_delta schema clamped
±1.0 → the instructed [-0.1, +0.3] band — an injected 'set my mastery
to 1.0' now fails validation into the #153 retry loop; plus a contract
test freezing that no tutor/note tool signature exposes
user_id/course_id/session_id/note_id to the model.
- Documented scope calls: the student's own message channel and session
history stay unwrapped (their instruction channel; wrapping the tool
but not message_history would be theater); catalog chunks stay trusted
(script-ingested official data); misconceptions tool stays
unregistered on the tutor (consent enforcement still deferred — the
data that DOES reach prompts is now contained); document-pipeline
workers untouched (no tools to coerce, ~70 cassettes at stake) — noted
as follow-up.
- Evals: chat_tutor (16) + quiz_generation (10) honestly re-recorded
(their prompts changed; replay keys on case names and would have stayed
silently green). The re-record also landed ADR-0023's tracked
prompt-shape fix (never end the turn on a tool call) after the quirk
reproduced live. Scores: all evaluators 1.000 on both datasets
(GroundedConcept ratcheted 0.875 → 1.0); other four datasets untouched.
27-case red-first injection suite (14 red at base) incl. an end-to-end
FunctionModel test proving an injected tool return reaches the model
enveloped. Gates: backend 1518 passed + ruff clean; lockvenv 278 passed;
evals replay green ×6.
Closes#150.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 30, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc09d245Commit Preview URL

Branch Preview URL
Jul 30 2026, 12:16 PM

@coderabbitai

coderabbitaiBot commented Jul 30, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:3 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 3ec87618-2e69-4198-9d09-789016b77862

📥 Commits

Reviewing files that changed from the base of the PR and between 21ac322 and c09d245.

📒 Files selected for processing (46)
  • backend/agents/chat_tutor.py
  • backend/agents/note_chat.py
  • backend/agents/note_concepts.py
  • backend/agents/note_summary.py
  • backend/agents/quiz.py
  • backend/agents/tools/chat_context.py
  • backend/agents/tools/graph.py
  • backend/agents/tools/graph_read.py
  • backend/agents/tools/note_context.py
  • backend/agents/tools/quiz_history.py
  • backend/prompts/preamble.txt
  • backend/routes/learn.py
  • backend/routes/notes.py
  • backend/services/graph_context.py
  • backend/services/prompt_safety.py
  • backend/services/rag_service.py
  • backend/tests/evals/baselines.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_big_o.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_dependency_injection.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_kantian_ethics.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_photosynthesis.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_supply_demand.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_chemistry_balancing.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_history_themes.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_intro_calculus.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_open_followup.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_python_recursion.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_stale_concept_review.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_advanced.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_correct_concept.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_minimal.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_misconception.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_partial_correct.json
  • backend/tests/evals/cassettes/quiz_generation/adaptive_downshift_struggling_student.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_basic_biology.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_intro_calculus.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_physics_definitions.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_real_analysis.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_theory_of_computation.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_data_structures.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_econ.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_organic_chem.json
  • backend/tests/evals/cassettes/quiz_generation/spaced_repetition_revives_stale_concept.json
  • backend/tests/test_graph_context_block.py
  • backend/tests/test_prompt_injection.py
  • docs/decisions/0023-tutor-graph-retrieval-seam.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Jul 30, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

…ero-width forgeries stripped, docs squared
- read_session_history_tool neutralizes replayed content (a student could
plant a literal envelope delimiter in one turn and have it handed back
as a REAL byte-match later); read_concepts_for_user_tool and
read_graph_neighborhood_tool neutralize student-derived concept names
(the same data the seed block already defangs). Three red-style tests.
- neutralize_delimiters strips zero-width/invisible Unicode before
matching — a visually identical forged delimiter threaded with U+200B
no longer dodges the regex (empirically demonstrated in review).
- Docs squared: ADR 0023's follow-up note marked shipped (via #470+#471),
prompt_safety's 'every agent' overclaim corrected (note workers carry
their own one-line guards, deliberately), graph_context's budget
docstring notes the envelope overhead.
- FYI-class same-PR fix: {last_session_summary} (LLM-generated text of
student content) now wrapped like the quiz digest.
Backend 1521 + ruff green; lockvenv 78 green; evals replay green ×6.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review pass complete: findings all fixed — the convergent one (sibling tools on the same agent returning raw content: session history could replay a student-planted REAL forged delimiter; both graph readers returned raw concept names) now neutralized with tests; zero-width delimiter forgeries stripped before matching; three doc staleness items squared; and the reviewers' FYI (last_session_summary unwrapped) folded in since it's the same class. send_to_tutor's preface deliberately stays unwrapped (the student's own instruction channel, consistent with the PR's documented scope call). Backend 1521 + lockvenv + evals ×6 green. e2e cycle next.

def test_session_history_tool_neutralizes_replayed_delimiters():
"""A student can plant a literal envelope delimiter in one turn; the
history tool must not replay it as a REAL byte-match later."""
import asyncio
def test_graph_read_tools_neutralize_concept_names():
"""The sibling read surfaces return the SAME student-derived concept
names the seed block neutralizes — they must be defanged here too."""
import asyncio
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Pre-merge gate: full lane 28/28 passed + oracles clean. Merging.

@AndresL230
AndresL230 merged commit 7703e22 into mainJul 30, 2026
8 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…#480)
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/b7-150-injection-hardening branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[P2] Agent migration: prompt-injection hardening on student-supplied content

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(agents): prompt-injection hardening on student content (#150) - #471

Merged
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening
Jul 30, 2026
Merged

feat(agents): prompt-injection hardening on student content (#150)#471
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

What

Third link of the B7 seam-endgame chain; gates the public beta. Full detail in the commit message. The shape:

  • Containment, single source of truth: services/prompt_safety.py — delimited untrusted-content envelopes with data-not-instructions framing and delimiter-forgery neutralization (embedded BEGIN/END copies are defanged, case/whitespace-insensitively, idempotently). Applied at every assembly boundary where student or peer-derived text enters a prompt: RAG chunks, the graph seed block, the legacy COURSE MATERIALS + shared-context JSON, and the tool→LLM boundary. Storage stays raw.
  • Instruction hardening: a shared injection-guard paragraph across the tutor's three preambles, note_chat, quiz, and the legacy preamble — plus the academic-integrity block held out of [P1] Agent migration: graph-grounded tutor retrieval + tool-use loop #149 for exactly this PR.
  • Tool-use constraint: the mastery-delta schema is clamped to the instructed band, so an injected "set my mastery to 1.0" fails validation into [P2] Agent platform: structured-output retry + validation hardening #153's retry loop; a contract test freezes that no tool signature exposes ids to the model.
  • Honest eval handling: replay keys on case names and would have stayed silently green on prompt changes — chat_tutor and quiz_generation were re-recorded (the re-record also landed ADR-0023's tracked prompt-shape fix after the bare-newline quirk reproduced live). All evaluators 1.000 on both datasets; GroundedConcept ratcheted 0.875 → 1.0.
  • Documented scope calls: the student's own message channel stays unwrapped (it is their instruction channel), catalog chunks stay trusted, the misconceptions tool stays unregistered pending real consent enforcement, and the tool-less document workers are a noted follow-up.

Verification

27-case red-first injection suite (14 red at base) including an end-to-end FunctionModel test proving injected tool returns reach the model enveloped. Backend 1518 passed + ruff clean; 278 passed under the lock-pinned 1.107 venv; evals replay green across all six datasets. Full local e2e cycle pre-merge; results below.

Closes#150.

🤖 Generated with Claude Code

Third link of the seam endgame; gates the public beta.
- One shared containment helper (services/prompt_safety.py):
wrap_untrusted builds a delimited BEGIN/END UNTRUSTED CONTENT envelope
with data-not-instructions framing; neutralize_delimiters defangs
embedded marker forgeries (case/whitespace-insensitive, idempotent) so
content can't fake an early END and escape. Applied at ASSEMBLY
boundaries only — storage keeps raw text: RAG chunks
(format_rag_context — covers agent chat, legacy chat, quiz context in
one place), the graph seed block's student-derived concept names, the
legacy prompt's COURSE MATERIALS + shared-context JSON, and the tool→
LLM boundary (search_course_materials, read_active_note,
read_misconceptions, quiz-history; note-worker user prompts).
- INJECTION_GUARD_PROMPT (single source) in the tutor's three preambles,
note_chat, quiz, and the legacy preamble; the ACADEMIC INTEGRITY block
deferred out of #149 lands in the agent preamble for parity.
- Tool-use constraint: ConceptMasteryUpdate.mastery_delta schema clamped
±1.0 → the instructed [-0.1, +0.3] band — an injected 'set my mastery
to 1.0' now fails validation into the #153 retry loop; plus a contract
test freezing that no tutor/note tool signature exposes
user_id/course_id/session_id/note_id to the model.
- Documented scope calls: the student's own message channel and session
history stay unwrapped (their instruction channel; wrapping the tool
but not message_history would be theater); catalog chunks stay trusted
(script-ingested official data); misconceptions tool stays
unregistered on the tutor (consent enforcement still deferred — the
data that DOES reach prompts is now contained); document-pipeline
workers untouched (no tools to coerce, ~70 cassettes at stake) — noted
as follow-up.
- Evals: chat_tutor (16) + quiz_generation (10) honestly re-recorded
(their prompts changed; replay keys on case names and would have stayed
silently green). The re-record also landed ADR-0023's tracked
prompt-shape fix (never end the turn on a tool call) after the quirk
reproduced live. Scores: all evaluators 1.000 on both datasets
(GroundedConcept ratcheted 0.875 → 1.0); other four datasets untouched.
27-case red-first injection suite (14 red at base) incl. an end-to-end
FunctionModel test proving an injected tool return reaches the model
enveloped. Gates: backend 1518 passed + ruff clean; lockvenv 278 passed;
evals replay green ×6.
Closes#150.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 30, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc09d245Commit Preview URL

Branch Preview URL
Jul 30 2026, 12:16 PM

@coderabbitai

coderabbitaiBot commented Jul 30, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:3 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 3ec87618-2e69-4198-9d09-789016b77862

📥 Commits

Reviewing files that changed from the base of the PR and between 21ac322 and c09d245.

📒 Files selected for processing (46)
  • backend/agents/chat_tutor.py
  • backend/agents/note_chat.py
  • backend/agents/note_concepts.py
  • backend/agents/note_summary.py
  • backend/agents/quiz.py
  • backend/agents/tools/chat_context.py
  • backend/agents/tools/graph.py
  • backend/agents/tools/graph_read.py
  • backend/agents/tools/note_context.py
  • backend/agents/tools/quiz_history.py
  • backend/prompts/preamble.txt
  • backend/routes/learn.py
  • backend/routes/notes.py
  • backend/services/graph_context.py
  • backend/services/prompt_safety.py
  • backend/services/rag_service.py
  • backend/tests/evals/baselines.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_big_o.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_dependency_injection.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_kantian_ethics.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_photosynthesis.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_supply_demand.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_chemistry_balancing.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_history_themes.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_intro_calculus.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_open_followup.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_python_recursion.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_stale_concept_review.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_advanced.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_correct_concept.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_minimal.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_misconception.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_partial_correct.json
  • backend/tests/evals/cassettes/quiz_generation/adaptive_downshift_struggling_student.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_basic_biology.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_intro_calculus.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_physics_definitions.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_real_analysis.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_theory_of_computation.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_data_structures.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_econ.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_organic_chem.json
  • backend/tests/evals/cassettes/quiz_generation/spaced_repetition_revives_stale_concept.json
  • backend/tests/test_graph_context_block.py
  • backend/tests/test_prompt_injection.py
  • docs/decisions/0023-tutor-graph-retrieval-seam.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Jul 30, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

…ero-width forgeries stripped, docs squared
- read_session_history_tool neutralizes replayed content (a student could
plant a literal envelope delimiter in one turn and have it handed back
as a REAL byte-match later); read_concepts_for_user_tool and
read_graph_neighborhood_tool neutralize student-derived concept names
(the same data the seed block already defangs). Three red-style tests.
- neutralize_delimiters strips zero-width/invisible Unicode before
matching — a visually identical forged delimiter threaded with U+200B
no longer dodges the regex (empirically demonstrated in review).
- Docs squared: ADR 0023's follow-up note marked shipped (via #470+#471),
prompt_safety's 'every agent' overclaim corrected (note workers carry
their own one-line guards, deliberately), graph_context's budget
docstring notes the envelope overhead.
- FYI-class same-PR fix: {last_session_summary} (LLM-generated text of
student content) now wrapped like the quiz digest.
Backend 1521 + ruff green; lockvenv 78 green; evals replay green ×6.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review pass complete: findings all fixed — the convergent one (sibling tools on the same agent returning raw content: session history could replay a student-planted REAL forged delimiter; both graph readers returned raw concept names) now neutralized with tests; zero-width delimiter forgeries stripped before matching; three doc staleness items squared; and the reviewers' FYI (last_session_summary unwrapped) folded in since it's the same class. send_to_tutor's preface deliberately stays unwrapped (the student's own instruction channel, consistent with the PR's documented scope call). Backend 1521 + lockvenv + evals ×6 green. e2e cycle next.

def test_session_history_tool_neutralizes_replayed_delimiters():
"""A student can plant a literal envelope delimiter in one turn; the
history tool must not replay it as a REAL byte-match later."""
import asyncio
def test_graph_read_tools_neutralize_concept_names():
"""The sibling read surfaces return the SAME student-derived concept
names the seed block neutralizes — they must be defanged here too."""
import asyncio
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Pre-merge gate: full lane 28/28 passed + oracles clean. Merging.

@AndresL230
AndresL230 merged commit 7703e22 into mainJul 30, 2026
8 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…#480)
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/b7-150-injection-hardening branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[P2] Agent migration: prompt-injection hardening on student-supplied content

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(agents): prompt-injection hardening on student content (#150) - #471

Merged
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening
Jul 30, 2026
Merged

feat(agents): prompt-injection hardening on student content (#150)#471
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

What

Third link of the B7 seam-endgame chain; gates the public beta. Full detail in the commit message. The shape:

  • Containment, single source of truth: services/prompt_safety.py — delimited untrusted-content envelopes with data-not-instructions framing and delimiter-forgery neutralization (embedded BEGIN/END copies are defanged, case/whitespace-insensitively, idempotently). Applied at every assembly boundary where student or peer-derived text enters a prompt: RAG chunks, the graph seed block, the legacy COURSE MATERIALS + shared-context JSON, and the tool→LLM boundary. Storage stays raw.
  • Instruction hardening: a shared injection-guard paragraph across the tutor's three preambles, note_chat, quiz, and the legacy preamble — plus the academic-integrity block held out of [P1] Agent migration: graph-grounded tutor retrieval + tool-use loop #149 for exactly this PR.
  • Tool-use constraint: the mastery-delta schema is clamped to the instructed band, so an injected "set my mastery to 1.0" fails validation into [P2] Agent platform: structured-output retry + validation hardening #153's retry loop; a contract test freezes that no tool signature exposes ids to the model.
  • Honest eval handling: replay keys on case names and would have stayed silently green on prompt changes — chat_tutor and quiz_generation were re-recorded (the re-record also landed ADR-0023's tracked prompt-shape fix after the bare-newline quirk reproduced live). All evaluators 1.000 on both datasets; GroundedConcept ratcheted 0.875 → 1.0.
  • Documented scope calls: the student's own message channel stays unwrapped (it is their instruction channel), catalog chunks stay trusted, the misconceptions tool stays unregistered pending real consent enforcement, and the tool-less document workers are a noted follow-up.

Verification

27-case red-first injection suite (14 red at base) including an end-to-end FunctionModel test proving injected tool returns reach the model enveloped. Backend 1518 passed + ruff clean; 278 passed under the lock-pinned 1.107 venv; evals replay green across all six datasets. Full local e2e cycle pre-merge; results below.

Closes#150.

🤖 Generated with Claude Code

Third link of the seam endgame; gates the public beta.
- One shared containment helper (services/prompt_safety.py):
wrap_untrusted builds a delimited BEGIN/END UNTRUSTED CONTENT envelope
with data-not-instructions framing; neutralize_delimiters defangs
embedded marker forgeries (case/whitespace-insensitive, idempotent) so
content can't fake an early END and escape. Applied at ASSEMBLY
boundaries only — storage keeps raw text: RAG chunks
(format_rag_context — covers agent chat, legacy chat, quiz context in
one place), the graph seed block's student-derived concept names, the
legacy prompt's COURSE MATERIALS + shared-context JSON, and the tool→
LLM boundary (search_course_materials, read_active_note,
read_misconceptions, quiz-history; note-worker user prompts).
- INJECTION_GUARD_PROMPT (single source) in the tutor's three preambles,
note_chat, quiz, and the legacy preamble; the ACADEMIC INTEGRITY block
deferred out of #149 lands in the agent preamble for parity.
- Tool-use constraint: ConceptMasteryUpdate.mastery_delta schema clamped
±1.0 → the instructed [-0.1, +0.3] band — an injected 'set my mastery
to 1.0' now fails validation into the #153 retry loop; plus a contract
test freezing that no tutor/note tool signature exposes
user_id/course_id/session_id/note_id to the model.
- Documented scope calls: the student's own message channel and session
history stay unwrapped (their instruction channel; wrapping the tool
but not message_history would be theater); catalog chunks stay trusted
(script-ingested official data); misconceptions tool stays
unregistered on the tutor (consent enforcement still deferred — the
data that DOES reach prompts is now contained); document-pipeline
workers untouched (no tools to coerce, ~70 cassettes at stake) — noted
as follow-up.
- Evals: chat_tutor (16) + quiz_generation (10) honestly re-recorded
(their prompts changed; replay keys on case names and would have stayed
silently green). The re-record also landed ADR-0023's tracked
prompt-shape fix (never end the turn on a tool call) after the quirk
reproduced live. Scores: all evaluators 1.000 on both datasets
(GroundedConcept ratcheted 0.875 → 1.0); other four datasets untouched.
27-case red-first injection suite (14 red at base) incl. an end-to-end
FunctionModel test proving an injected tool return reaches the model
enveloped. Gates: backend 1518 passed + ruff clean; lockvenv 278 passed;
evals replay green ×6.
Closes#150.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 30, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc09d245Commit Preview URL

Branch Preview URL
Jul 30 2026, 12:16 PM

@coderabbitai

coderabbitaiBot commented Jul 30, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:3 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 3ec87618-2e69-4198-9d09-789016b77862

📥 Commits

Reviewing files that changed from the base of the PR and between 21ac322 and c09d245.

📒 Files selected for processing (46)
  • backend/agents/chat_tutor.py
  • backend/agents/note_chat.py
  • backend/agents/note_concepts.py
  • backend/agents/note_summary.py
  • backend/agents/quiz.py
  • backend/agents/tools/chat_context.py
  • backend/agents/tools/graph.py
  • backend/agents/tools/graph_read.py
  • backend/agents/tools/note_context.py
  • backend/agents/tools/quiz_history.py
  • backend/prompts/preamble.txt
  • backend/routes/learn.py
  • backend/routes/notes.py
  • backend/services/graph_context.py
  • backend/services/prompt_safety.py
  • backend/services/rag_service.py
  • backend/tests/evals/baselines.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_big_o.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_dependency_injection.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_kantian_ethics.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_photosynthesis.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_supply_demand.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_chemistry_balancing.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_history_themes.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_intro_calculus.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_open_followup.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_python_recursion.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_stale_concept_review.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_advanced.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_correct_concept.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_minimal.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_misconception.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_partial_correct.json
  • backend/tests/evals/cassettes/quiz_generation/adaptive_downshift_struggling_student.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_basic_biology.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_intro_calculus.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_physics_definitions.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_real_analysis.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_theory_of_computation.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_data_structures.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_econ.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_organic_chem.json
  • backend/tests/evals/cassettes/quiz_generation/spaced_repetition_revives_stale_concept.json
  • backend/tests/test_graph_context_block.py
  • backend/tests/test_prompt_injection.py
  • docs/decisions/0023-tutor-graph-retrieval-seam.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Jul 30, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

…ero-width forgeries stripped, docs squared
- read_session_history_tool neutralizes replayed content (a student could
plant a literal envelope delimiter in one turn and have it handed back
as a REAL byte-match later); read_concepts_for_user_tool and
read_graph_neighborhood_tool neutralize student-derived concept names
(the same data the seed block already defangs). Three red-style tests.
- neutralize_delimiters strips zero-width/invisible Unicode before
matching — a visually identical forged delimiter threaded with U+200B
no longer dodges the regex (empirically demonstrated in review).
- Docs squared: ADR 0023's follow-up note marked shipped (via #470+#471),
prompt_safety's 'every agent' overclaim corrected (note workers carry
their own one-line guards, deliberately), graph_context's budget
docstring notes the envelope overhead.
- FYI-class same-PR fix: {last_session_summary} (LLM-generated text of
student content) now wrapped like the quiz digest.
Backend 1521 + ruff green; lockvenv 78 green; evals replay green ×6.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review pass complete: findings all fixed — the convergent one (sibling tools on the same agent returning raw content: session history could replay a student-planted REAL forged delimiter; both graph readers returned raw concept names) now neutralized with tests; zero-width delimiter forgeries stripped before matching; three doc staleness items squared; and the reviewers' FYI (last_session_summary unwrapped) folded in since it's the same class. send_to_tutor's preface deliberately stays unwrapped (the student's own instruction channel, consistent with the PR's documented scope call). Backend 1521 + lockvenv + evals ×6 green. e2e cycle next.

def test_session_history_tool_neutralizes_replayed_delimiters():
"""A student can plant a literal envelope delimiter in one turn; the
history tool must not replay it as a REAL byte-match later."""
import asyncio
def test_graph_read_tools_neutralize_concept_names():
"""The sibling read surfaces return the SAME student-derived concept
names the seed block neutralizes — they must be defanged here too."""
import asyncio
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Pre-merge gate: full lane 28/28 passed + oracles clean. Merging.

@AndresL230
AndresL230 merged commit 7703e22 into mainJul 30, 2026
8 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…#480)
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/b7-150-injection-hardening branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[P2] Agent migration: prompt-injection hardening on student-supplied content

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(agents): prompt-injection hardening on student content (#150) - #471

Merged
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening
Jul 30, 2026
Merged

feat(agents): prompt-injection hardening on student content (#150)#471
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

What

Third link of the B7 seam-endgame chain; gates the public beta. Full detail in the commit message. The shape:

  • Containment, single source of truth: services/prompt_safety.py — delimited untrusted-content envelopes with data-not-instructions framing and delimiter-forgery neutralization (embedded BEGIN/END copies are defanged, case/whitespace-insensitively, idempotently). Applied at every assembly boundary where student or peer-derived text enters a prompt: RAG chunks, the graph seed block, the legacy COURSE MATERIALS + shared-context JSON, and the tool→LLM boundary. Storage stays raw.
  • Instruction hardening: a shared injection-guard paragraph across the tutor's three preambles, note_chat, quiz, and the legacy preamble — plus the academic-integrity block held out of [P1] Agent migration: graph-grounded tutor retrieval + tool-use loop #149 for exactly this PR.
  • Tool-use constraint: the mastery-delta schema is clamped to the instructed band, so an injected "set my mastery to 1.0" fails validation into [P2] Agent platform: structured-output retry + validation hardening #153's retry loop; a contract test freezes that no tool signature exposes ids to the model.
  • Honest eval handling: replay keys on case names and would have stayed silently green on prompt changes — chat_tutor and quiz_generation were re-recorded (the re-record also landed ADR-0023's tracked prompt-shape fix after the bare-newline quirk reproduced live). All evaluators 1.000 on both datasets; GroundedConcept ratcheted 0.875 → 1.0.
  • Documented scope calls: the student's own message channel stays unwrapped (it is their instruction channel), catalog chunks stay trusted, the misconceptions tool stays unregistered pending real consent enforcement, and the tool-less document workers are a noted follow-up.

Verification

27-case red-first injection suite (14 red at base) including an end-to-end FunctionModel test proving injected tool returns reach the model enveloped. Backend 1518 passed + ruff clean; 278 passed under the lock-pinned 1.107 venv; evals replay green across all six datasets. Full local e2e cycle pre-merge; results below.

Closes#150.

🤖 Generated with Claude Code

Third link of the seam endgame; gates the public beta.
- One shared containment helper (services/prompt_safety.py):
wrap_untrusted builds a delimited BEGIN/END UNTRUSTED CONTENT envelope
with data-not-instructions framing; neutralize_delimiters defangs
embedded marker forgeries (case/whitespace-insensitive, idempotent) so
content can't fake an early END and escape. Applied at ASSEMBLY
boundaries only — storage keeps raw text: RAG chunks
(format_rag_context — covers agent chat, legacy chat, quiz context in
one place), the graph seed block's student-derived concept names, the
legacy prompt's COURSE MATERIALS + shared-context JSON, and the tool→
LLM boundary (search_course_materials, read_active_note,
read_misconceptions, quiz-history; note-worker user prompts).
- INJECTION_GUARD_PROMPT (single source) in the tutor's three preambles,
note_chat, quiz, and the legacy preamble; the ACADEMIC INTEGRITY block
deferred out of #149 lands in the agent preamble for parity.
- Tool-use constraint: ConceptMasteryUpdate.mastery_delta schema clamped
±1.0 → the instructed [-0.1, +0.3] band — an injected 'set my mastery
to 1.0' now fails validation into the #153 retry loop; plus a contract
test freezing that no tutor/note tool signature exposes
user_id/course_id/session_id/note_id to the model.
- Documented scope calls: the student's own message channel and session
history stay unwrapped (their instruction channel; wrapping the tool
but not message_history would be theater); catalog chunks stay trusted
(script-ingested official data); misconceptions tool stays
unregistered on the tutor (consent enforcement still deferred — the
data that DOES reach prompts is now contained); document-pipeline
workers untouched (no tools to coerce, ~70 cassettes at stake) — noted
as follow-up.
- Evals: chat_tutor (16) + quiz_generation (10) honestly re-recorded
(their prompts changed; replay keys on case names and would have stayed
silently green). The re-record also landed ADR-0023's tracked
prompt-shape fix (never end the turn on a tool call) after the quirk
reproduced live. Scores: all evaluators 1.000 on both datasets
(GroundedConcept ratcheted 0.875 → 1.0); other four datasets untouched.
27-case red-first injection suite (14 red at base) incl. an end-to-end
FunctionModel test proving an injected tool return reaches the model
enveloped. Gates: backend 1518 passed + ruff clean; lockvenv 278 passed;
evals replay green ×6.
Closes#150.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 30, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc09d245Commit Preview URL

Branch Preview URL
Jul 30 2026, 12:16 PM

@coderabbitai

coderabbitaiBot commented Jul 30, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:3 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 3ec87618-2e69-4198-9d09-789016b77862

📥 Commits

Reviewing files that changed from the base of the PR and between 21ac322 and c09d245.

📒 Files selected for processing (46)
  • backend/agents/chat_tutor.py
  • backend/agents/note_chat.py
  • backend/agents/note_concepts.py
  • backend/agents/note_summary.py
  • backend/agents/quiz.py
  • backend/agents/tools/chat_context.py
  • backend/agents/tools/graph.py
  • backend/agents/tools/graph_read.py
  • backend/agents/tools/note_context.py
  • backend/agents/tools/quiz_history.py
  • backend/prompts/preamble.txt
  • backend/routes/learn.py
  • backend/routes/notes.py
  • backend/services/graph_context.py
  • backend/services/prompt_safety.py
  • backend/services/rag_service.py
  • backend/tests/evals/baselines.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_big_o.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_dependency_injection.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_kantian_ethics.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_photosynthesis.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_supply_demand.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_chemistry_balancing.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_history_themes.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_intro_calculus.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_open_followup.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_python_recursion.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_stale_concept_review.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_advanced.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_correct_concept.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_minimal.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_misconception.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_partial_correct.json
  • backend/tests/evals/cassettes/quiz_generation/adaptive_downshift_struggling_student.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_basic_biology.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_intro_calculus.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_physics_definitions.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_real_analysis.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_theory_of_computation.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_data_structures.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_econ.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_organic_chem.json
  • backend/tests/evals/cassettes/quiz_generation/spaced_repetition_revives_stale_concept.json
  • backend/tests/test_graph_context_block.py
  • backend/tests/test_prompt_injection.py
  • docs/decisions/0023-tutor-graph-retrieval-seam.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Jul 30, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

…ero-width forgeries stripped, docs squared
- read_session_history_tool neutralizes replayed content (a student could
plant a literal envelope delimiter in one turn and have it handed back
as a REAL byte-match later); read_concepts_for_user_tool and
read_graph_neighborhood_tool neutralize student-derived concept names
(the same data the seed block already defangs). Three red-style tests.
- neutralize_delimiters strips zero-width/invisible Unicode before
matching — a visually identical forged delimiter threaded with U+200B
no longer dodges the regex (empirically demonstrated in review).
- Docs squared: ADR 0023's follow-up note marked shipped (via #470+#471),
prompt_safety's 'every agent' overclaim corrected (note workers carry
their own one-line guards, deliberately), graph_context's budget
docstring notes the envelope overhead.
- FYI-class same-PR fix: {last_session_summary} (LLM-generated text of
student content) now wrapped like the quiz digest.
Backend 1521 + ruff green; lockvenv 78 green; evals replay green ×6.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review pass complete: findings all fixed — the convergent one (sibling tools on the same agent returning raw content: session history could replay a student-planted REAL forged delimiter; both graph readers returned raw concept names) now neutralized with tests; zero-width delimiter forgeries stripped before matching; three doc staleness items squared; and the reviewers' FYI (last_session_summary unwrapped) folded in since it's the same class. send_to_tutor's preface deliberately stays unwrapped (the student's own instruction channel, consistent with the PR's documented scope call). Backend 1521 + lockvenv + evals ×6 green. e2e cycle next.

def test_session_history_tool_neutralizes_replayed_delimiters():
"""A student can plant a literal envelope delimiter in one turn; the
history tool must not replay it as a REAL byte-match later."""
import asyncio
def test_graph_read_tools_neutralize_concept_names():
"""The sibling read surfaces return the SAME student-derived concept
names the seed block neutralizes — they must be defanged here too."""
import asyncio
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Pre-merge gate: full lane 28/28 passed + oracles clean. Merging.

@AndresL230
AndresL230 merged commit 7703e22 into mainJul 30, 2026
8 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…#480)
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/b7-150-injection-hardening branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[P2] Agent migration: prompt-injection hardening on student-supplied content

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(agents): prompt-injection hardening on student content (#150) - #471

Merged
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening
Jul 30, 2026
Merged

feat(agents): prompt-injection hardening on student content (#150)#471
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

What

Third link of the B7 seam-endgame chain; gates the public beta. Full detail in the commit message. The shape:

  • Containment, single source of truth: services/prompt_safety.py — delimited untrusted-content envelopes with data-not-instructions framing and delimiter-forgery neutralization (embedded BEGIN/END copies are defanged, case/whitespace-insensitively, idempotently). Applied at every assembly boundary where student or peer-derived text enters a prompt: RAG chunks, the graph seed block, the legacy COURSE MATERIALS + shared-context JSON, and the tool→LLM boundary. Storage stays raw.
  • Instruction hardening: a shared injection-guard paragraph across the tutor's three preambles, note_chat, quiz, and the legacy preamble — plus the academic-integrity block held out of [P1] Agent migration: graph-grounded tutor retrieval + tool-use loop #149 for exactly this PR.
  • Tool-use constraint: the mastery-delta schema is clamped to the instructed band, so an injected "set my mastery to 1.0" fails validation into [P2] Agent platform: structured-output retry + validation hardening #153's retry loop; a contract test freezes that no tool signature exposes ids to the model.
  • Honest eval handling: replay keys on case names and would have stayed silently green on prompt changes — chat_tutor and quiz_generation were re-recorded (the re-record also landed ADR-0023's tracked prompt-shape fix after the bare-newline quirk reproduced live). All evaluators 1.000 on both datasets; GroundedConcept ratcheted 0.875 → 1.0.
  • Documented scope calls: the student's own message channel stays unwrapped (it is their instruction channel), catalog chunks stay trusted, the misconceptions tool stays unregistered pending real consent enforcement, and the tool-less document workers are a noted follow-up.

Verification

27-case red-first injection suite (14 red at base) including an end-to-end FunctionModel test proving injected tool returns reach the model enveloped. Backend 1518 passed + ruff clean; 278 passed under the lock-pinned 1.107 venv; evals replay green across all six datasets. Full local e2e cycle pre-merge; results below.

Closes#150.

🤖 Generated with Claude Code

Third link of the seam endgame; gates the public beta.
- One shared containment helper (services/prompt_safety.py):
wrap_untrusted builds a delimited BEGIN/END UNTRUSTED CONTENT envelope
with data-not-instructions framing; neutralize_delimiters defangs
embedded marker forgeries (case/whitespace-insensitive, idempotent) so
content can't fake an early END and escape. Applied at ASSEMBLY
boundaries only — storage keeps raw text: RAG chunks
(format_rag_context — covers agent chat, legacy chat, quiz context in
one place), the graph seed block's student-derived concept names, the
legacy prompt's COURSE MATERIALS + shared-context JSON, and the tool→
LLM boundary (search_course_materials, read_active_note,
read_misconceptions, quiz-history; note-worker user prompts).
- INJECTION_GUARD_PROMPT (single source) in the tutor's three preambles,
note_chat, quiz, and the legacy preamble; the ACADEMIC INTEGRITY block
deferred out of #149 lands in the agent preamble for parity.
- Tool-use constraint: ConceptMasteryUpdate.mastery_delta schema clamped
±1.0 → the instructed [-0.1, +0.3] band — an injected 'set my mastery
to 1.0' now fails validation into the #153 retry loop; plus a contract
test freezing that no tutor/note tool signature exposes
user_id/course_id/session_id/note_id to the model.
- Documented scope calls: the student's own message channel and session
history stay unwrapped (their instruction channel; wrapping the tool
but not message_history would be theater); catalog chunks stay trusted
(script-ingested official data); misconceptions tool stays
unregistered on the tutor (consent enforcement still deferred — the
data that DOES reach prompts is now contained); document-pipeline
workers untouched (no tools to coerce, ~70 cassettes at stake) — noted
as follow-up.
- Evals: chat_tutor (16) + quiz_generation (10) honestly re-recorded
(their prompts changed; replay keys on case names and would have stayed
silently green). The re-record also landed ADR-0023's tracked
prompt-shape fix (never end the turn on a tool call) after the quirk
reproduced live. Scores: all evaluators 1.000 on both datasets
(GroundedConcept ratcheted 0.875 → 1.0); other four datasets untouched.
27-case red-first injection suite (14 red at base) incl. an end-to-end
FunctionModel test proving an injected tool return reaches the model
enveloped. Gates: backend 1518 passed + ruff clean; lockvenv 278 passed;
evals replay green ×6.
Closes#150.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 30, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc09d245Commit Preview URL

Branch Preview URL
Jul 30 2026, 12:16 PM

@coderabbitai

coderabbitaiBot commented Jul 30, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:3 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 3ec87618-2e69-4198-9d09-789016b77862

📥 Commits

Reviewing files that changed from the base of the PR and between 21ac322 and c09d245.

📒 Files selected for processing (46)
  • backend/agents/chat_tutor.py
  • backend/agents/note_chat.py
  • backend/agents/note_concepts.py
  • backend/agents/note_summary.py
  • backend/agents/quiz.py
  • backend/agents/tools/chat_context.py
  • backend/agents/tools/graph.py
  • backend/agents/tools/graph_read.py
  • backend/agents/tools/note_context.py
  • backend/agents/tools/quiz_history.py
  • backend/prompts/preamble.txt
  • backend/routes/learn.py
  • backend/routes/notes.py
  • backend/services/graph_context.py
  • backend/services/prompt_safety.py
  • backend/services/rag_service.py
  • backend/tests/evals/baselines.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_big_o.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_dependency_injection.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_kantian_ethics.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_photosynthesis.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_supply_demand.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_chemistry_balancing.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_history_themes.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_intro_calculus.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_open_followup.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_python_recursion.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_stale_concept_review.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_advanced.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_correct_concept.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_minimal.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_misconception.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_partial_correct.json
  • backend/tests/evals/cassettes/quiz_generation/adaptive_downshift_struggling_student.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_basic_biology.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_intro_calculus.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_physics_definitions.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_real_analysis.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_theory_of_computation.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_data_structures.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_econ.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_organic_chem.json
  • backend/tests/evals/cassettes/quiz_generation/spaced_repetition_revives_stale_concept.json
  • backend/tests/test_graph_context_block.py
  • backend/tests/test_prompt_injection.py
  • docs/decisions/0023-tutor-graph-retrieval-seam.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Jul 30, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

…ero-width forgeries stripped, docs squared
- read_session_history_tool neutralizes replayed content (a student could
plant a literal envelope delimiter in one turn and have it handed back
as a REAL byte-match later); read_concepts_for_user_tool and
read_graph_neighborhood_tool neutralize student-derived concept names
(the same data the seed block already defangs). Three red-style tests.
- neutralize_delimiters strips zero-width/invisible Unicode before
matching — a visually identical forged delimiter threaded with U+200B
no longer dodges the regex (empirically demonstrated in review).
- Docs squared: ADR 0023's follow-up note marked shipped (via #470+#471),
prompt_safety's 'every agent' overclaim corrected (note workers carry
their own one-line guards, deliberately), graph_context's budget
docstring notes the envelope overhead.
- FYI-class same-PR fix: {last_session_summary} (LLM-generated text of
student content) now wrapped like the quiz digest.
Backend 1521 + ruff green; lockvenv 78 green; evals replay green ×6.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review pass complete: findings all fixed — the convergent one (sibling tools on the same agent returning raw content: session history could replay a student-planted REAL forged delimiter; both graph readers returned raw concept names) now neutralized with tests; zero-width delimiter forgeries stripped before matching; three doc staleness items squared; and the reviewers' FYI (last_session_summary unwrapped) folded in since it's the same class. send_to_tutor's preface deliberately stays unwrapped (the student's own instruction channel, consistent with the PR's documented scope call). Backend 1521 + lockvenv + evals ×6 green. e2e cycle next.

def test_session_history_tool_neutralizes_replayed_delimiters():
"""A student can plant a literal envelope delimiter in one turn; the
history tool must not replay it as a REAL byte-match later."""
import asyncio
def test_graph_read_tools_neutralize_concept_names():
"""The sibling read surfaces return the SAME student-derived concept
names the seed block neutralizes — they must be defanged here too."""
import asyncio
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Pre-merge gate: full lane 28/28 passed + oracles clean. Merging.

@AndresL230
AndresL230 merged commit 7703e22 into mainJul 30, 2026
8 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…#480)
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/b7-150-injection-hardening branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[P2] Agent migration: prompt-injection hardening on student-supplied content

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(agents): prompt-injection hardening on student content (#150) - #471

Merged
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening
Jul 30, 2026
Merged

feat(agents): prompt-injection hardening on student content (#150)#471
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

What

Third link of the B7 seam-endgame chain; gates the public beta. Full detail in the commit message. The shape:

  • Containment, single source of truth: services/prompt_safety.py — delimited untrusted-content envelopes with data-not-instructions framing and delimiter-forgery neutralization (embedded BEGIN/END copies are defanged, case/whitespace-insensitively, idempotently). Applied at every assembly boundary where student or peer-derived text enters a prompt: RAG chunks, the graph seed block, the legacy COURSE MATERIALS + shared-context JSON, and the tool→LLM boundary. Storage stays raw.
  • Instruction hardening: a shared injection-guard paragraph across the tutor's three preambles, note_chat, quiz, and the legacy preamble — plus the academic-integrity block held out of [P1] Agent migration: graph-grounded tutor retrieval + tool-use loop #149 for exactly this PR.
  • Tool-use constraint: the mastery-delta schema is clamped to the instructed band, so an injected "set my mastery to 1.0" fails validation into [P2] Agent platform: structured-output retry + validation hardening #153's retry loop; a contract test freezes that no tool signature exposes ids to the model.
  • Honest eval handling: replay keys on case names and would have stayed silently green on prompt changes — chat_tutor and quiz_generation were re-recorded (the re-record also landed ADR-0023's tracked prompt-shape fix after the bare-newline quirk reproduced live). All evaluators 1.000 on both datasets; GroundedConcept ratcheted 0.875 → 1.0.
  • Documented scope calls: the student's own message channel stays unwrapped (it is their instruction channel), catalog chunks stay trusted, the misconceptions tool stays unregistered pending real consent enforcement, and the tool-less document workers are a noted follow-up.

Verification

27-case red-first injection suite (14 red at base) including an end-to-end FunctionModel test proving injected tool returns reach the model enveloped. Backend 1518 passed + ruff clean; 278 passed under the lock-pinned 1.107 venv; evals replay green across all six datasets. Full local e2e cycle pre-merge; results below.

Closes#150.

🤖 Generated with Claude Code

Third link of the seam endgame; gates the public beta.
- One shared containment helper (services/prompt_safety.py):
wrap_untrusted builds a delimited BEGIN/END UNTRUSTED CONTENT envelope
with data-not-instructions framing; neutralize_delimiters defangs
embedded marker forgeries (case/whitespace-insensitive, idempotent) so
content can't fake an early END and escape. Applied at ASSEMBLY
boundaries only — storage keeps raw text: RAG chunks
(format_rag_context — covers agent chat, legacy chat, quiz context in
one place), the graph seed block's student-derived concept names, the
legacy prompt's COURSE MATERIALS + shared-context JSON, and the tool→
LLM boundary (search_course_materials, read_active_note,
read_misconceptions, quiz-history; note-worker user prompts).
- INJECTION_GUARD_PROMPT (single source) in the tutor's three preambles,
note_chat, quiz, and the legacy preamble; the ACADEMIC INTEGRITY block
deferred out of #149 lands in the agent preamble for parity.
- Tool-use constraint: ConceptMasteryUpdate.mastery_delta schema clamped
±1.0 → the instructed [-0.1, +0.3] band — an injected 'set my mastery
to 1.0' now fails validation into the #153 retry loop; plus a contract
test freezing that no tutor/note tool signature exposes
user_id/course_id/session_id/note_id to the model.
- Documented scope calls: the student's own message channel and session
history stay unwrapped (their instruction channel; wrapping the tool
but not message_history would be theater); catalog chunks stay trusted
(script-ingested official data); misconceptions tool stays
unregistered on the tutor (consent enforcement still deferred — the
data that DOES reach prompts is now contained); document-pipeline
workers untouched (no tools to coerce, ~70 cassettes at stake) — noted
as follow-up.
- Evals: chat_tutor (16) + quiz_generation (10) honestly re-recorded
(their prompts changed; replay keys on case names and would have stayed
silently green). The re-record also landed ADR-0023's tracked
prompt-shape fix (never end the turn on a tool call) after the quirk
reproduced live. Scores: all evaluators 1.000 on both datasets
(GroundedConcept ratcheted 0.875 → 1.0); other four datasets untouched.
27-case red-first injection suite (14 red at base) incl. an end-to-end
FunctionModel test proving an injected tool return reaches the model
enveloped. Gates: backend 1518 passed + ruff clean; lockvenv 278 passed;
evals replay green ×6.
Closes#150.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 30, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc09d245Commit Preview URL

Branch Preview URL
Jul 30 2026, 12:16 PM

@coderabbitai

coderabbitaiBot commented Jul 30, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:3 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 3ec87618-2e69-4198-9d09-789016b77862

📥 Commits

Reviewing files that changed from the base of the PR and between 21ac322 and c09d245.

📒 Files selected for processing (46)
  • backend/agents/chat_tutor.py
  • backend/agents/note_chat.py
  • backend/agents/note_concepts.py
  • backend/agents/note_summary.py
  • backend/agents/quiz.py
  • backend/agents/tools/chat_context.py
  • backend/agents/tools/graph.py
  • backend/agents/tools/graph_read.py
  • backend/agents/tools/note_context.py
  • backend/agents/tools/quiz_history.py
  • backend/prompts/preamble.txt
  • backend/routes/learn.py
  • backend/routes/notes.py
  • backend/services/graph_context.py
  • backend/services/prompt_safety.py
  • backend/services/rag_service.py
  • backend/tests/evals/baselines.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_big_o.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_dependency_injection.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_kantian_ethics.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_photosynthesis.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_supply_demand.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_chemistry_balancing.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_history_themes.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_intro_calculus.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_open_followup.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_python_recursion.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_stale_concept_review.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_advanced.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_correct_concept.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_minimal.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_misconception.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_partial_correct.json
  • backend/tests/evals/cassettes/quiz_generation/adaptive_downshift_struggling_student.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_basic_biology.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_intro_calculus.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_physics_definitions.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_real_analysis.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_theory_of_computation.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_data_structures.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_econ.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_organic_chem.json
  • backend/tests/evals/cassettes/quiz_generation/spaced_repetition_revives_stale_concept.json
  • backend/tests/test_graph_context_block.py
  • backend/tests/test_prompt_injection.py
  • docs/decisions/0023-tutor-graph-retrieval-seam.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Jul 30, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

…ero-width forgeries stripped, docs squared
- read_session_history_tool neutralizes replayed content (a student could
plant a literal envelope delimiter in one turn and have it handed back
as a REAL byte-match later); read_concepts_for_user_tool and
read_graph_neighborhood_tool neutralize student-derived concept names
(the same data the seed block already defangs). Three red-style tests.
- neutralize_delimiters strips zero-width/invisible Unicode before
matching — a visually identical forged delimiter threaded with U+200B
no longer dodges the regex (empirically demonstrated in review).
- Docs squared: ADR 0023's follow-up note marked shipped (via #470+#471),
prompt_safety's 'every agent' overclaim corrected (note workers carry
their own one-line guards, deliberately), graph_context's budget
docstring notes the envelope overhead.
- FYI-class same-PR fix: {last_session_summary} (LLM-generated text of
student content) now wrapped like the quiz digest.
Backend 1521 + ruff green; lockvenv 78 green; evals replay green ×6.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review pass complete: findings all fixed — the convergent one (sibling tools on the same agent returning raw content: session history could replay a student-planted REAL forged delimiter; both graph readers returned raw concept names) now neutralized with tests; zero-width delimiter forgeries stripped before matching; three doc staleness items squared; and the reviewers' FYI (last_session_summary unwrapped) folded in since it's the same class. send_to_tutor's preface deliberately stays unwrapped (the student's own instruction channel, consistent with the PR's documented scope call). Backend 1521 + lockvenv + evals ×6 green. e2e cycle next.

def test_session_history_tool_neutralizes_replayed_delimiters():
"""A student can plant a literal envelope delimiter in one turn; the
history tool must not replay it as a REAL byte-match later."""
import asyncio
def test_graph_read_tools_neutralize_concept_names():
"""The sibling read surfaces return the SAME student-derived concept
names the seed block neutralizes — they must be defanged here too."""
import asyncio
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Pre-merge gate: full lane 28/28 passed + oracles clean. Merging.

@AndresL230
AndresL230 merged commit 7703e22 into mainJul 30, 2026
8 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…#480)
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/b7-150-injection-hardening branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[P2] Agent migration: prompt-injection hardening on student-supplied content

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(agents): prompt-injection hardening on student content (#150) - #471

Merged
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening
Jul 30, 2026
Merged

feat(agents): prompt-injection hardening on student content (#150)#471
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

What

Third link of the B7 seam-endgame chain; gates the public beta. Full detail in the commit message. The shape:

  • Containment, single source of truth: services/prompt_safety.py — delimited untrusted-content envelopes with data-not-instructions framing and delimiter-forgery neutralization (embedded BEGIN/END copies are defanged, case/whitespace-insensitively, idempotently). Applied at every assembly boundary where student or peer-derived text enters a prompt: RAG chunks, the graph seed block, the legacy COURSE MATERIALS + shared-context JSON, and the tool→LLM boundary. Storage stays raw.
  • Instruction hardening: a shared injection-guard paragraph across the tutor's three preambles, note_chat, quiz, and the legacy preamble — plus the academic-integrity block held out of [P1] Agent migration: graph-grounded tutor retrieval + tool-use loop #149 for exactly this PR.
  • Tool-use constraint: the mastery-delta schema is clamped to the instructed band, so an injected "set my mastery to 1.0" fails validation into [P2] Agent platform: structured-output retry + validation hardening #153's retry loop; a contract test freezes that no tool signature exposes ids to the model.
  • Honest eval handling: replay keys on case names and would have stayed silently green on prompt changes — chat_tutor and quiz_generation were re-recorded (the re-record also landed ADR-0023's tracked prompt-shape fix after the bare-newline quirk reproduced live). All evaluators 1.000 on both datasets; GroundedConcept ratcheted 0.875 → 1.0.
  • Documented scope calls: the student's own message channel stays unwrapped (it is their instruction channel), catalog chunks stay trusted, the misconceptions tool stays unregistered pending real consent enforcement, and the tool-less document workers are a noted follow-up.

Verification

27-case red-first injection suite (14 red at base) including an end-to-end FunctionModel test proving injected tool returns reach the model enveloped. Backend 1518 passed + ruff clean; 278 passed under the lock-pinned 1.107 venv; evals replay green across all six datasets. Full local e2e cycle pre-merge; results below.

Closes#150.

🤖 Generated with Claude Code

Third link of the seam endgame; gates the public beta.
- One shared containment helper (services/prompt_safety.py):
wrap_untrusted builds a delimited BEGIN/END UNTRUSTED CONTENT envelope
with data-not-instructions framing; neutralize_delimiters defangs
embedded marker forgeries (case/whitespace-insensitive, idempotent) so
content can't fake an early END and escape. Applied at ASSEMBLY
boundaries only — storage keeps raw text: RAG chunks
(format_rag_context — covers agent chat, legacy chat, quiz context in
one place), the graph seed block's student-derived concept names, the
legacy prompt's COURSE MATERIALS + shared-context JSON, and the tool→
LLM boundary (search_course_materials, read_active_note,
read_misconceptions, quiz-history; note-worker user prompts).
- INJECTION_GUARD_PROMPT (single source) in the tutor's three preambles,
note_chat, quiz, and the legacy preamble; the ACADEMIC INTEGRITY block
deferred out of #149 lands in the agent preamble for parity.
- Tool-use constraint: ConceptMasteryUpdate.mastery_delta schema clamped
±1.0 → the instructed [-0.1, +0.3] band — an injected 'set my mastery
to 1.0' now fails validation into the #153 retry loop; plus a contract
test freezing that no tutor/note tool signature exposes
user_id/course_id/session_id/note_id to the model.
- Documented scope calls: the student's own message channel and session
history stay unwrapped (their instruction channel; wrapping the tool
but not message_history would be theater); catalog chunks stay trusted
(script-ingested official data); misconceptions tool stays
unregistered on the tutor (consent enforcement still deferred — the
data that DOES reach prompts is now contained); document-pipeline
workers untouched (no tools to coerce, ~70 cassettes at stake) — noted
as follow-up.
- Evals: chat_tutor (16) + quiz_generation (10) honestly re-recorded
(their prompts changed; replay keys on case names and would have stayed
silently green). The re-record also landed ADR-0023's tracked
prompt-shape fix (never end the turn on a tool call) after the quirk
reproduced live. Scores: all evaluators 1.000 on both datasets
(GroundedConcept ratcheted 0.875 → 1.0); other four datasets untouched.
27-case red-first injection suite (14 red at base) incl. an end-to-end
FunctionModel test proving an injected tool return reaches the model
enveloped. Gates: backend 1518 passed + ruff clean; lockvenv 278 passed;
evals replay green ×6.
Closes#150.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 30, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc09d245Commit Preview URL

Branch Preview URL
Jul 30 2026, 12:16 PM

@coderabbitai

coderabbitaiBot commented Jul 30, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:3 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 3ec87618-2e69-4198-9d09-789016b77862

📥 Commits

Reviewing files that changed from the base of the PR and between 21ac322 and c09d245.

📒 Files selected for processing (46)
  • backend/agents/chat_tutor.py
  • backend/agents/note_chat.py
  • backend/agents/note_concepts.py
  • backend/agents/note_summary.py
  • backend/agents/quiz.py
  • backend/agents/tools/chat_context.py
  • backend/agents/tools/graph.py
  • backend/agents/tools/graph_read.py
  • backend/agents/tools/note_context.py
  • backend/agents/tools/quiz_history.py
  • backend/prompts/preamble.txt
  • backend/routes/learn.py
  • backend/routes/notes.py
  • backend/services/graph_context.py
  • backend/services/prompt_safety.py
  • backend/services/rag_service.py
  • backend/tests/evals/baselines.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_big_o.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_dependency_injection.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_kantian_ethics.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_photosynthesis.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_supply_demand.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_chemistry_balancing.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_history_themes.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_intro_calculus.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_open_followup.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_python_recursion.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_stale_concept_review.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_advanced.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_correct_concept.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_minimal.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_misconception.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_partial_correct.json
  • backend/tests/evals/cassettes/quiz_generation/adaptive_downshift_struggling_student.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_basic_biology.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_intro_calculus.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_physics_definitions.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_real_analysis.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_theory_of_computation.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_data_structures.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_econ.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_organic_chem.json
  • backend/tests/evals/cassettes/quiz_generation/spaced_repetition_revives_stale_concept.json
  • backend/tests/test_graph_context_block.py
  • backend/tests/test_prompt_injection.py
  • docs/decisions/0023-tutor-graph-retrieval-seam.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Jul 30, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

…ero-width forgeries stripped, docs squared
- read_session_history_tool neutralizes replayed content (a student could
plant a literal envelope delimiter in one turn and have it handed back
as a REAL byte-match later); read_concepts_for_user_tool and
read_graph_neighborhood_tool neutralize student-derived concept names
(the same data the seed block already defangs). Three red-style tests.
- neutralize_delimiters strips zero-width/invisible Unicode before
matching — a visually identical forged delimiter threaded with U+200B
no longer dodges the regex (empirically demonstrated in review).
- Docs squared: ADR 0023's follow-up note marked shipped (via #470+#471),
prompt_safety's 'every agent' overclaim corrected (note workers carry
their own one-line guards, deliberately), graph_context's budget
docstring notes the envelope overhead.
- FYI-class same-PR fix: {last_session_summary} (LLM-generated text of
student content) now wrapped like the quiz digest.
Backend 1521 + ruff green; lockvenv 78 green; evals replay green ×6.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review pass complete: findings all fixed — the convergent one (sibling tools on the same agent returning raw content: session history could replay a student-planted REAL forged delimiter; both graph readers returned raw concept names) now neutralized with tests; zero-width delimiter forgeries stripped before matching; three doc staleness items squared; and the reviewers' FYI (last_session_summary unwrapped) folded in since it's the same class. send_to_tutor's preface deliberately stays unwrapped (the student's own instruction channel, consistent with the PR's documented scope call). Backend 1521 + lockvenv + evals ×6 green. e2e cycle next.

def test_session_history_tool_neutralizes_replayed_delimiters():
"""A student can plant a literal envelope delimiter in one turn; the
history tool must not replay it as a REAL byte-match later."""
import asyncio
def test_graph_read_tools_neutralize_concept_names():
"""The sibling read surfaces return the SAME student-derived concept
names the seed block neutralizes — they must be defanged here too."""
import asyncio
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Pre-merge gate: full lane 28/28 passed + oracles clean. Merging.

@AndresL230
AndresL230 merged commit 7703e22 into mainJul 30, 2026
8 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…#480)
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/b7-150-injection-hardening branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[P2] Agent migration: prompt-injection hardening on student-supplied content

1 participant

@AndresL230
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(agents): prompt-injection hardening on student content (#150) - #471

Merged
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening
Jul 30, 2026
Merged

feat(agents): prompt-injection hardening on student content (#150)#471
AndresL230 merged 2 commits into
mainfrom
feat/b7-150-injection-hardening

Conversation

@AndresL230

Copy link
Copy Markdown
Collaborator

What

Third link of the B7 seam-endgame chain; gates the public beta. Full detail in the commit message. The shape:

  • Containment, single source of truth: services/prompt_safety.py — delimited untrusted-content envelopes with data-not-instructions framing and delimiter-forgery neutralization (embedded BEGIN/END copies are defanged, case/whitespace-insensitively, idempotently). Applied at every assembly boundary where student or peer-derived text enters a prompt: RAG chunks, the graph seed block, the legacy COURSE MATERIALS + shared-context JSON, and the tool→LLM boundary. Storage stays raw.
  • Instruction hardening: a shared injection-guard paragraph across the tutor's three preambles, note_chat, quiz, and the legacy preamble — plus the academic-integrity block held out of [P1] Agent migration: graph-grounded tutor retrieval + tool-use loop #149 for exactly this PR.
  • Tool-use constraint: the mastery-delta schema is clamped to the instructed band, so an injected "set my mastery to 1.0" fails validation into [P2] Agent platform: structured-output retry + validation hardening #153's retry loop; a contract test freezes that no tool signature exposes ids to the model.
  • Honest eval handling: replay keys on case names and would have stayed silently green on prompt changes — chat_tutor and quiz_generation were re-recorded (the re-record also landed ADR-0023's tracked prompt-shape fix after the bare-newline quirk reproduced live). All evaluators 1.000 on both datasets; GroundedConcept ratcheted 0.875 → 1.0.
  • Documented scope calls: the student's own message channel stays unwrapped (it is their instruction channel), catalog chunks stay trusted, the misconceptions tool stays unregistered pending real consent enforcement, and the tool-less document workers are a noted follow-up.

Verification

27-case red-first injection suite (14 red at base) including an end-to-end FunctionModel test proving injected tool returns reach the model enveloped. Backend 1518 passed + ruff clean; 278 passed under the lock-pinned 1.107 venv; evals replay green across all six datasets. Full local e2e cycle pre-merge; results below.

Closes#150.

🤖 Generated with Claude Code

Third link of the seam endgame; gates the public beta.
- One shared containment helper (services/prompt_safety.py):
wrap_untrusted builds a delimited BEGIN/END UNTRUSTED CONTENT envelope
with data-not-instructions framing; neutralize_delimiters defangs
embedded marker forgeries (case/whitespace-insensitive, idempotent) so
content can't fake an early END and escape. Applied at ASSEMBLY
boundaries only — storage keeps raw text: RAG chunks
(format_rag_context — covers agent chat, legacy chat, quiz context in
one place), the graph seed block's student-derived concept names, the
legacy prompt's COURSE MATERIALS + shared-context JSON, and the tool→
LLM boundary (search_course_materials, read_active_note,
read_misconceptions, quiz-history; note-worker user prompts).
- INJECTION_GUARD_PROMPT (single source) in the tutor's three preambles,
note_chat, quiz, and the legacy preamble; the ACADEMIC INTEGRITY block
deferred out of #149 lands in the agent preamble for parity.
- Tool-use constraint: ConceptMasteryUpdate.mastery_delta schema clamped
±1.0 → the instructed [-0.1, +0.3] band — an injected 'set my mastery
to 1.0' now fails validation into the #153 retry loop; plus a contract
test freezing that no tutor/note tool signature exposes
user_id/course_id/session_id/note_id to the model.
- Documented scope calls: the student's own message channel and session
history stay unwrapped (their instruction channel; wrapping the tool
but not message_history would be theater); catalog chunks stay trusted
(script-ingested official data); misconceptions tool stays
unregistered on the tutor (consent enforcement still deferred — the
data that DOES reach prompts is now contained); document-pipeline
workers untouched (no tools to coerce, ~70 cassettes at stake) — noted
as follow-up.
- Evals: chat_tutor (16) + quiz_generation (10) honestly re-recorded
(their prompts changed; replay keys on case names and would have stayed
silently green). The re-record also landed ADR-0023's tracked
prompt-shape fix (never end the turn on a tool call) after the quirk
reproduced live. Scores: all evaluators 1.000 on both datasets
(GroundedConcept ratcheted 0.875 → 1.0); other four datasets untouched.
27-case red-first injection suite (14 red at base) incl. an end-to-end
FunctionModel test proving an injected tool return reaches the model
enveloped. Gates: backend 1518 passed + ruff clean; lockvenv 278 passed;
evals replay green ×6.
Closes#150.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pagesBot commented Jul 30, 2026

Copy link
Copy Markdown

Deploying with Cloudflare Workers Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

StatusNameLatest CommitPreview URLUpdated (UTC)
✅ Deployment successful!
View logs
frontend-stagingc09d245Commit Preview URL

Branch Preview URL
Jul 30 2026, 12:16 PM

@coderabbitai

coderabbitaiBot commented Jul 30, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AndresL230, you've reached your PR review limit, so we couldn't start this review.

Next review available in:3 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 3ec87618-2e69-4198-9d09-789016b77862

📥 Commits

Reviewing files that changed from the base of the PR and between 21ac322 and c09d245.

📒 Files selected for processing (46)
  • backend/agents/chat_tutor.py
  • backend/agents/note_chat.py
  • backend/agents/note_concepts.py
  • backend/agents/note_summary.py
  • backend/agents/quiz.py
  • backend/agents/tools/chat_context.py
  • backend/agents/tools/graph.py
  • backend/agents/tools/graph_read.py
  • backend/agents/tools/note_context.py
  • backend/agents/tools/quiz_history.py
  • backend/prompts/preamble.txt
  • backend/routes/learn.py
  • backend/routes/notes.py
  • backend/services/graph_context.py
  • backend/services/prompt_safety.py
  • backend/services/rag_service.py
  • backend/tests/evals/baselines.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_big_o.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_dependency_injection.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_kantian_ethics.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_photosynthesis.json
  • backend/tests/evals/cassettes/chat_tutor/expository_explain_supply_demand.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_chemistry_balancing.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_history_themes.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_intro_calculus.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_open_followup.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_python_recursion.json
  • backend/tests/evals/cassettes/chat_tutor/socratic_stale_concept_review.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_advanced.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_correct_concept.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_minimal.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_misconception.json
  • backend/tests/evals/cassettes/chat_tutor/teachback_partial_correct.json
  • backend/tests/evals/cassettes/quiz_generation/adaptive_downshift_struggling_student.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_basic_biology.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_intro_calculus.json
  • backend/tests/evals/cassettes/quiz_generation/easy_mcq_physics_definitions.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_real_analysis.json
  • backend/tests/evals/cassettes/quiz_generation/hard_mcq_theory_of_computation.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_data_structures.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_econ.json
  • backend/tests/evals/cassettes/quiz_generation/medium_mcq_organic_chem.json
  • backend/tests/evals/cassettes/quiz_generation/spaced_repetition_revives_stale_concept.json
  • backend/tests/test_graph_context_block.py
  • backend/tests/test_prompt_injection.py
  • docs/decisions/0023-tutor-graph-retrieval-seam.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@supabase

supabaseBot commented Jul 30, 2026

Copy link
Copy Markdown

This pull request has been ignored for the connected project ybgqdonkoqftwrmweuyv because there are no changes detected in supabase directory. You can change this behaviour in Project Integrations Settings ↗︎.


Preview Branches by Supabase.
Learn more about Supabase Branching ↗︎.

…ero-width forgeries stripped, docs squared
- read_session_history_tool neutralizes replayed content (a student could
plant a literal envelope delimiter in one turn and have it handed back
as a REAL byte-match later); read_concepts_for_user_tool and
read_graph_neighborhood_tool neutralize student-derived concept names
(the same data the seed block already defangs). Three red-style tests.
- neutralize_delimiters strips zero-width/invisible Unicode before
matching — a visually identical forged delimiter threaded with U+200B
no longer dodges the regex (empirically demonstrated in review).
- Docs squared: ADR 0023's follow-up note marked shipped (via #470+#471),
prompt_safety's 'every agent' overclaim corrected (note workers carry
their own one-line guards, deliberately), graph_context's budget
docstring notes the envelope overhead.
- FYI-class same-PR fix: {last_session_summary} (LLM-generated text of
student content) now wrapped like the quiz digest.
Backend 1521 + ruff green; lockvenv 78 green; evals replay green ×6.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Review pass complete: findings all fixed — the convergent one (sibling tools on the same agent returning raw content: session history could replay a student-planted REAL forged delimiter; both graph readers returned raw concept names) now neutralized with tests; zero-width delimiter forgeries stripped before matching; three doc staleness items squared; and the reviewers' FYI (last_session_summary unwrapped) folded in since it's the same class. send_to_tutor's preface deliberately stays unwrapped (the student's own instruction channel, consistent with the PR's documented scope call). Backend 1521 + lockvenv + evals ×6 green. e2e cycle next.

def test_session_history_tool_neutralizes_replayed_delimiters():
"""A student can plant a literal envelope delimiter in one turn; the
history tool must not replay it as a REAL byte-match later."""
import asyncio
def test_graph_read_tools_neutralize_concept_names():
"""The sibling read surfaces return the SAME student-derived concept
names the seed block neutralizes — they must be defanged here too."""
import asyncio
@AndresL230

Copy link
Copy Markdown
CollaboratorAuthor

Pre-merge gate: full lane 28/28 passed + oracles clean. Merging.

@AndresL230
AndresL230 merged commit 7703e22 into mainJul 30, 2026
8 checks passed
AndresL230 added a commit that referenced this pull request Jul 30, 2026
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AndresL230 added a commit that referenced this pull request Jul 30, 2026
…#480)
Every CI run since the #471 merge fails at the lint step with exit 2:
the agent-only rung ladder cutover removed three suppressed
no-restricted-syntax occurrences from screens/Learn.tsx, and eslint
exits 2 when suppressions in eslint-suppressions.json no longer match
anything. Prune the stale entries (20 -> 17) via --prune-suppressions.
Verified against origin/main: eslint exit 0, tsc exit 0, vitest 372/372.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@AndresL230
AndresL230 deleted the feat/b7-150-injection-hardening branch August 2, 2026 18:30
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[P2] Agent migration: prompt-injection hardening on student-supplied content

1 participant

@AndresL230