Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -34,6 +34,9 @@ dist/
*.sqlite
*.sqlite3

# ---- RAG data (large files — stored in Supabase Storage) ----
backend/data/*.json

# ---- Env / secrets ----
.env
.env.*
Expand Down
210 changes: 210 additions & 0 deletions backend/data/scrape_summary.json
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,210 @@
{
"total_courses": 8720,
"total_errors": 50,
"errors": [
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-650a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-or-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-529a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-641a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-525a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-511a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-541a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-532a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-660a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-523a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-644a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-581a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pa-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-534a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-md-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-546a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-542a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-519a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-527a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-544a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-540a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-532a/",
"error": "empty response"
}
],
Comment on lines +2 to +205

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Don't commit a scrape summary that already records 50 failed course pages.

Ingestion uses this scrape output as the source corpus, so these unrecovered SDM URLs mean the initial RAG index is knowingly incomplete. Please either rerun until this is clean or make ingestion fail fast when the scrape summary reports unresolved errors.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` around lines 2 - 205, The scrape summary
currently includes unresolved failures, so do not keep this snapshot as-is.
Update the ingestion flow that reads scrape_summary.json to either rerun
scraping until the errors array is empty or make the ingest step fail fast when
total_errors is nonzero, using the summary fields total_errors and errors as the
check. Locate the validation in the scrape-summary ingestion path and ensure it
blocks indexing when any failed course URLs remain.

"elapsed_seconds": 1890,
"completed_at": "2026-06-27T01:39:55.726124+00:00",
"semester_tag": "fall_2026",
"output_file": "C:\\Users\\Jack\\Desktop\\VS Code\\sapling\\backend\\data\\bu_catalog_fall_2026.json"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

Remove the machine-local absolute path from this committed artifact.

C:\Users\Jack\... leaks workstation details and makes the summary non-portable. A repo-relative path or no path field at all would avoid noisy diffs and the local identifier exposure.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` at line 209, The committed scrape_summary
artifact currently hardcodes a machine-local absolute path in the output_file
field, which leaks workstation-specific details and breaks portability. Update
the value in scrape_summary.json to use a repo-relative path or remove the field
entirely, and ensure any code that writes this summary uses a stable path source
so future runs do not reintroduce local identifiers; check the summary
generation logic and the output_file serialization path.

}
8 changes: 8 additions & 0 deletions backend/db/connection.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -101,3 +101,11 @@ def delete(self, filters: dict) -> list:

def table(name: str) -> SupabaseTable:
return SupabaseTable(name)


def rpc(function_name: str, params: dict) -> list:
"""Call a Supabase Postgres function via /rest/v1/rpc/{function_name}."""
url = f"{REST_URL}/rpc/{function_name}"
r = _client.post(url, json=params)
r.raise_for_status()
return r.json()
119 changes: 119 additions & 0 deletions backend/routes/documents.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -785,6 +785,7 @@ async def event_stream():
("invalidate_study_guide_cache", _invalidate_study_guide_cache, user_id, course_id),
("update_course_context", update_course_context, course_id),
("check_upload_achievements", _check_upload_achievements, user_id),
("index_document_chunks", _index_document_chunks, doc_id, course_id, user_id, extracted_text, classification.category, getattr(summary, "abstract", "")),
)

yield sapling_event_to_sse(SaplingEvent(
Expand DownExpand Up@@ -898,6 +899,124 @@ def _check_upload_achievements(user_id: str) -> None:
pass


def _chunk_text(text: str, chunk_size: int = 800, overlap: int = 100) -> list[str]:
"""Split text into overlapping character-window chunks."""
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end].strip())
start += chunk_size - overlap
return [c for c in chunks if len(c) > 50] # drop near-empty tail chunks


def _index_document_chunks(
doc_id: str,
course_id: str, # Sapling UUID — resolved to BU code internally
user_id: str,
extracted_text: str,
category: str,
doc_summary: str = "",
) -> None:
"""Chunk, embed, and upsert a document into course_chunks.

Runs in a background thread via _spawn_post_roll after the document
is persisted, so it never blocks the SSE stream.
"""
import hashlib
import math
from google import genai as _genai
from google.genai import types as genai_types
from db.connection import table
import os, time

MIN_COURSE_RELEVANCE = 0.35 # below this, document is likely off-topic for the course

try:
# Resolve BU course code from Sapling UUID
rows = table("courses").select(
"course_code", filters={"id": f"eq.{course_id}"}, limit=1
)
bu_course_id = (rows[0].get("course_code") or course_id) if rows else course_id

chunks = _chunk_text(extracted_text)
if not chunks:
return

_gclient = _genai.Client(api_key=os.getenv("GEMINI_API_KEY", ""))

def _embed_texts(texts: list[str]) -> list[list[float]]:
resp = _gclient.models.embed_content(
model="gemini-embedding-001",
contents=texts,
config=genai_types.EmbedContentConfig(output_dimensionality=768),
)
return [list(e.values) for e in resp.embeddings]

# ── Relevance gate ────────────────────────────────────────────────────
# Fetch the catalog chunk embedding for this course and compare against
# the document's first chunk. Irrelevant documents are skipped to keep
# the index clean.
catalog_rows = table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
if catalog_rows and catalog_rows[0].get("embedding"):
catalog_vec = catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))
if dot < MIN_COURSE_RELEVANCE:
Comment on lines +960 to +975

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Compare against the best catalog match, not an arbitrary catalog row.

Line 960 fetches limit=1 with no ordering, so the relevance gate can reject an on-topic upload if Supabase returns a non-representative catalog chunk. Fetch the course’s catalog embeddings and gate on the max cosine score instead.

Suggested adjustment
- catalog_rows = table("course_chunks").select(- "embedding",- filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},- limit=1,- )- if catalog_rows and catalog_rows[0].get("embedding"):- catalog_vec = catalog_rows[0]["embedding"]+ catalog_rows = table("course_chunks").select(+ "embedding",+ filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},+ )+ catalog_vecs = [r["embedding"] for r in catalog_rows if r.get("embedding")]+ if catalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
- # cosine similarity (vectors are unit-norm from the model)- dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))+ dot = max(sum(a * b for a, b in zip(doc_sample_vec, catalog_vec)) for catalog_vec in catalog_vecs)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
ifcatalog_rowsandcatalog_rows[0].get("embedding"):
catalog_vec=catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot=sum(a*bfora, binzip(doc_sample_vec, catalog_vec))
ifdot<MIN_COURSE_RELEVANCE:
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
)
catalog_vecs= [r["embedding"] forrincatalog_rowsifr.get("embedding")]
ifcatalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
dot=max(sum(a*bfora, binzip(doc_sample_vec, catalog_vec)) forcatalog_vecincatalog_vecs)
ifdot<MIN_COURSE_RELEVANCE:
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 960 - 975, The relevance check
currently uses only the first returned catalog chunk from
table("course_chunks"), which is arbitrary because the query has limit=1 without
ordering. Update the logic around the catalog_rows lookup to evaluate all
catalog embeddings for the course and compare the document sample against the
best cosine match, then apply the MIN_COURSE_RELEVANCE gate to that maximum
score. Keep the change localized near the existing sample_text, doc_sample_vec,
and dot calculation so the representative catalog selection is no longer
dependent on Supabase row order.

logger.warning(
"[RAG] doc %s skipped — relevance to %s is %.3f (< %.2f)",
doc_id, bu_course_id, dot, MIN_COURSE_RELEVANCE,
)
return

records = []
for i, chunk_text in enumerate(chunks):
raw = f"{doc_id}::{i}::{chunk_text}"
cid = hashlib.sha256(raw.encode()).hexdigest()
records.append({
"id": cid,
"course_id": bu_course_id,
"doc_id": doc_id,
"uploader_id": user_id,
"chunk_index": i,
"chunk_text": chunk_text,
"chunk_hash": cid,
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})

# Embed in batches of 50
BATCH = 50
for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota

table("course_chunks").upsert(records, on_conflict="id")
Comment on lines +994 to +1014

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Only upsert successfully embedded chunks, and keep upserts batched.

If _embed_texts fails for a batch, those records keep embedding=None but are still sent to course_chunks; additionally, all chunks are posted in one large request after embedding. That can either pollute the retrieval index or make large uploads fail at the final Supabase call. Upsert only records with embeddings per batch, matching the catalog ingest pattern.

Suggested adjustment
 for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
+ if len(vecs) != len(batch):+ raise ValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
+ table("course_chunks").upsert(batch, on_conflict="id")
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
- table("course_chunks").upsert(records, on_conflict="id")
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
table("course_chunks").upsert(records, on_conflict="id")
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
iflen(vecs) !=len(batch):
raiseValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
table("course_chunks").upsert(batch, on_conflict="id")
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 994 - 1014, The batch embedding
flow in the document ingest path still upserts every record at the end,
including chunks whose embedding stayed None after _embed_texts failed. Update
the course_chunks ingest logic to mirror the catalog ingest pattern by upserting
only the successfully embedded records inside each batch loop, and skip failed
batch records instead of carrying them forward. Keep the batching behavior in
place around the existing record processing in the documents route so large
uploads do not get sent as one final upsert.

logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
except Exception:
logger.exception("[RAG] _index_document_chunks failed for doc %s", doc_id)


def _spawn_post_roll(*tasks: tuple) -> None:
"""Fire-and-forget post-roll work for SSE / non-FastAPI-BackgroundTasks
contexts. Each tuple is (label, callable, *args). Exceptions in the
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -34,6 +34,9 @@ dist/
*.sqlite
*.sqlite3

# ---- RAG data (large files — stored in Supabase Storage) ----
backend/data/*.json

# ---- Env / secrets ----
.env
.env.*
Expand Down
210 changes: 210 additions & 0 deletions backend/data/scrape_summary.json
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,210 @@
{
"total_courses": 8720,
"total_errors": 50,
"errors": [
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-650a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-or-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-529a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-641a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-525a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-511a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-541a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-532a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-660a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-523a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-644a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-581a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pa-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-534a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-md-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-546a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-542a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-519a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-527a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-544a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-540a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-532a/",
"error": "empty response"
}
],
Comment on lines +2 to +205

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Don't commit a scrape summary that already records 50 failed course pages.

Ingestion uses this scrape output as the source corpus, so these unrecovered SDM URLs mean the initial RAG index is knowingly incomplete. Please either rerun until this is clean or make ingestion fail fast when the scrape summary reports unresolved errors.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` around lines 2 - 205, The scrape summary
currently includes unresolved failures, so do not keep this snapshot as-is.
Update the ingestion flow that reads scrape_summary.json to either rerun
scraping until the errors array is empty or make the ingest step fail fast when
total_errors is nonzero, using the summary fields total_errors and errors as the
check. Locate the validation in the scrape-summary ingestion path and ensure it
blocks indexing when any failed course URLs remain.

"elapsed_seconds": 1890,
"completed_at": "2026-06-27T01:39:55.726124+00:00",
"semester_tag": "fall_2026",
"output_file": "C:\\Users\\Jack\\Desktop\\VS Code\\sapling\\backend\\data\\bu_catalog_fall_2026.json"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

Remove the machine-local absolute path from this committed artifact.

C:\Users\Jack\... leaks workstation details and makes the summary non-portable. A repo-relative path or no path field at all would avoid noisy diffs and the local identifier exposure.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` at line 209, The committed scrape_summary
artifact currently hardcodes a machine-local absolute path in the output_file
field, which leaks workstation-specific details and breaks portability. Update
the value in scrape_summary.json to use a repo-relative path or remove the field
entirely, and ensure any code that writes this summary uses a stable path source
so future runs do not reintroduce local identifiers; check the summary
generation logic and the output_file serialization path.

}
8 changes: 8 additions & 0 deletions backend/db/connection.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -101,3 +101,11 @@ def delete(self, filters: dict) -> list:

def table(name: str) -> SupabaseTable:
return SupabaseTable(name)


def rpc(function_name: str, params: dict) -> list:
"""Call a Supabase Postgres function via /rest/v1/rpc/{function_name}."""
url = f"{REST_URL}/rpc/{function_name}"
r = _client.post(url, json=params)
r.raise_for_status()
return r.json()
119 changes: 119 additions & 0 deletions backend/routes/documents.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -785,6 +785,7 @@ async def event_stream():
("invalidate_study_guide_cache", _invalidate_study_guide_cache, user_id, course_id),
("update_course_context", update_course_context, course_id),
("check_upload_achievements", _check_upload_achievements, user_id),
("index_document_chunks", _index_document_chunks, doc_id, course_id, user_id, extracted_text, classification.category, getattr(summary, "abstract", "")),
)

yield sapling_event_to_sse(SaplingEvent(
Expand DownExpand Up@@ -898,6 +899,124 @@ def _check_upload_achievements(user_id: str) -> None:
pass


def _chunk_text(text: str, chunk_size: int = 800, overlap: int = 100) -> list[str]:
"""Split text into overlapping character-window chunks."""
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end].strip())
start += chunk_size - overlap
return [c for c in chunks if len(c) > 50] # drop near-empty tail chunks


def _index_document_chunks(
doc_id: str,
course_id: str, # Sapling UUID — resolved to BU code internally
user_id: str,
extracted_text: str,
category: str,
doc_summary: str = "",
) -> None:
"""Chunk, embed, and upsert a document into course_chunks.

Runs in a background thread via _spawn_post_roll after the document
is persisted, so it never blocks the SSE stream.
"""
import hashlib
import math
from google import genai as _genai
from google.genai import types as genai_types
from db.connection import table
import os, time

MIN_COURSE_RELEVANCE = 0.35 # below this, document is likely off-topic for the course

try:
# Resolve BU course code from Sapling UUID
rows = table("courses").select(
"course_code", filters={"id": f"eq.{course_id}"}, limit=1
)
bu_course_id = (rows[0].get("course_code") or course_id) if rows else course_id

chunks = _chunk_text(extracted_text)
if not chunks:
return

_gclient = _genai.Client(api_key=os.getenv("GEMINI_API_KEY", ""))

def _embed_texts(texts: list[str]) -> list[list[float]]:
resp = _gclient.models.embed_content(
model="gemini-embedding-001",
contents=texts,
config=genai_types.EmbedContentConfig(output_dimensionality=768),
)
return [list(e.values) for e in resp.embeddings]

# ── Relevance gate ────────────────────────────────────────────────────
# Fetch the catalog chunk embedding for this course and compare against
# the document's first chunk. Irrelevant documents are skipped to keep
# the index clean.
catalog_rows = table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
if catalog_rows and catalog_rows[0].get("embedding"):
catalog_vec = catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))
if dot < MIN_COURSE_RELEVANCE:
Comment on lines +960 to +975

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Compare against the best catalog match, not an arbitrary catalog row.

Line 960 fetches limit=1 with no ordering, so the relevance gate can reject an on-topic upload if Supabase returns a non-representative catalog chunk. Fetch the course’s catalog embeddings and gate on the max cosine score instead.

Suggested adjustment
- catalog_rows = table("course_chunks").select(- "embedding",- filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},- limit=1,- )- if catalog_rows and catalog_rows[0].get("embedding"):- catalog_vec = catalog_rows[0]["embedding"]+ catalog_rows = table("course_chunks").select(+ "embedding",+ filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},+ )+ catalog_vecs = [r["embedding"] for r in catalog_rows if r.get("embedding")]+ if catalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
- # cosine similarity (vectors are unit-norm from the model)- dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))+ dot = max(sum(a * b for a, b in zip(doc_sample_vec, catalog_vec)) for catalog_vec in catalog_vecs)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
ifcatalog_rowsandcatalog_rows[0].get("embedding"):
catalog_vec=catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot=sum(a*bfora, binzip(doc_sample_vec, catalog_vec))
ifdot<MIN_COURSE_RELEVANCE:
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
)
catalog_vecs= [r["embedding"] forrincatalog_rowsifr.get("embedding")]
ifcatalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
dot=max(sum(a*bfora, binzip(doc_sample_vec, catalog_vec)) forcatalog_vecincatalog_vecs)
ifdot<MIN_COURSE_RELEVANCE:
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 960 - 975, The relevance check
currently uses only the first returned catalog chunk from
table("course_chunks"), which is arbitrary because the query has limit=1 without
ordering. Update the logic around the catalog_rows lookup to evaluate all
catalog embeddings for the course and compare the document sample against the
best cosine match, then apply the MIN_COURSE_RELEVANCE gate to that maximum
score. Keep the change localized near the existing sample_text, doc_sample_vec,
and dot calculation so the representative catalog selection is no longer
dependent on Supabase row order.

logger.warning(
"[RAG] doc %s skipped — relevance to %s is %.3f (< %.2f)",
doc_id, bu_course_id, dot, MIN_COURSE_RELEVANCE,
)
return

records = []
for i, chunk_text in enumerate(chunks):
raw = f"{doc_id}::{i}::{chunk_text}"
cid = hashlib.sha256(raw.encode()).hexdigest()
records.append({
"id": cid,
"course_id": bu_course_id,
"doc_id": doc_id,
"uploader_id": user_id,
"chunk_index": i,
"chunk_text": chunk_text,
"chunk_hash": cid,
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})

# Embed in batches of 50
BATCH = 50
for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota

table("course_chunks").upsert(records, on_conflict="id")
Comment on lines +994 to +1014

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Only upsert successfully embedded chunks, and keep upserts batched.

If _embed_texts fails for a batch, those records keep embedding=None but are still sent to course_chunks; additionally, all chunks are posted in one large request after embedding. That can either pollute the retrieval index or make large uploads fail at the final Supabase call. Upsert only records with embeddings per batch, matching the catalog ingest pattern.

Suggested adjustment
 for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
+ if len(vecs) != len(batch):+ raise ValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
+ table("course_chunks").upsert(batch, on_conflict="id")
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
- table("course_chunks").upsert(records, on_conflict="id")
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
table("course_chunks").upsert(records, on_conflict="id")
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
iflen(vecs) !=len(batch):
raiseValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
table("course_chunks").upsert(batch, on_conflict="id")
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 994 - 1014, The batch embedding
flow in the document ingest path still upserts every record at the end,
including chunks whose embedding stayed None after _embed_texts failed. Update
the course_chunks ingest logic to mirror the catalog ingest pattern by upserting
only the successfully embedded records inside each batch loop, and skip failed
batch records instead of carrying them forward. Keep the batching behavior in
place around the existing record processing in the documents route so large
uploads do not get sent as one final upsert.

logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
except Exception:
logger.exception("[RAG] _index_document_chunks failed for doc %s", doc_id)


def _spawn_post_roll(*tasks: tuple) -> None:
"""Fire-and-forget post-roll work for SSE / non-FastAPI-BackgroundTasks
contexts. Each tuple is (label, callable, *args). Exceptions in the
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -34,6 +34,9 @@ dist/
*.sqlite
*.sqlite3

# ---- RAG data (large files — stored in Supabase Storage) ----
backend/data/*.json

# ---- Env / secrets ----
.env
.env.*
Expand Down
210 changes: 210 additions & 0 deletions backend/data/scrape_summary.json
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,210 @@
{
"total_courses": 8720,
"total_errors": 50,
"errors": [
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-650a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-or-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-529a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-641a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-525a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-511a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-541a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-532a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-660a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-523a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-644a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-581a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pa-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-534a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-md-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-546a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-542a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-519a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-527a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-544a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-540a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-532a/",
"error": "empty response"
}
],
Comment on lines +2 to +205

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Don't commit a scrape summary that already records 50 failed course pages.

Ingestion uses this scrape output as the source corpus, so these unrecovered SDM URLs mean the initial RAG index is knowingly incomplete. Please either rerun until this is clean or make ingestion fail fast when the scrape summary reports unresolved errors.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` around lines 2 - 205, The scrape summary
currently includes unresolved failures, so do not keep this snapshot as-is.
Update the ingestion flow that reads scrape_summary.json to either rerun
scraping until the errors array is empty or make the ingest step fail fast when
total_errors is nonzero, using the summary fields total_errors and errors as the
check. Locate the validation in the scrape-summary ingestion path and ensure it
blocks indexing when any failed course URLs remain.

"elapsed_seconds": 1890,
"completed_at": "2026-06-27T01:39:55.726124+00:00",
"semester_tag": "fall_2026",
"output_file": "C:\\Users\\Jack\\Desktop\\VS Code\\sapling\\backend\\data\\bu_catalog_fall_2026.json"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

Remove the machine-local absolute path from this committed artifact.

C:\Users\Jack\... leaks workstation details and makes the summary non-portable. A repo-relative path or no path field at all would avoid noisy diffs and the local identifier exposure.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` at line 209, The committed scrape_summary
artifact currently hardcodes a machine-local absolute path in the output_file
field, which leaks workstation-specific details and breaks portability. Update
the value in scrape_summary.json to use a repo-relative path or remove the field
entirely, and ensure any code that writes this summary uses a stable path source
so future runs do not reintroduce local identifiers; check the summary
generation logic and the output_file serialization path.

}
8 changes: 8 additions & 0 deletions backend/db/connection.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -101,3 +101,11 @@ def delete(self, filters: dict) -> list:

def table(name: str) -> SupabaseTable:
return SupabaseTable(name)


def rpc(function_name: str, params: dict) -> list:
"""Call a Supabase Postgres function via /rest/v1/rpc/{function_name}."""
url = f"{REST_URL}/rpc/{function_name}"
r = _client.post(url, json=params)
r.raise_for_status()
return r.json()
119 changes: 119 additions & 0 deletions backend/routes/documents.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -785,6 +785,7 @@ async def event_stream():
("invalidate_study_guide_cache", _invalidate_study_guide_cache, user_id, course_id),
("update_course_context", update_course_context, course_id),
("check_upload_achievements", _check_upload_achievements, user_id),
("index_document_chunks", _index_document_chunks, doc_id, course_id, user_id, extracted_text, classification.category, getattr(summary, "abstract", "")),
)

yield sapling_event_to_sse(SaplingEvent(
Expand DownExpand Up@@ -898,6 +899,124 @@ def _check_upload_achievements(user_id: str) -> None:
pass


def _chunk_text(text: str, chunk_size: int = 800, overlap: int = 100) -> list[str]:
"""Split text into overlapping character-window chunks."""
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end].strip())
start += chunk_size - overlap
return [c for c in chunks if len(c) > 50] # drop near-empty tail chunks


def _index_document_chunks(
doc_id: str,
course_id: str, # Sapling UUID — resolved to BU code internally
user_id: str,
extracted_text: str,
category: str,
doc_summary: str = "",
) -> None:
"""Chunk, embed, and upsert a document into course_chunks.

Runs in a background thread via _spawn_post_roll after the document
is persisted, so it never blocks the SSE stream.
"""
import hashlib
import math
from google import genai as _genai
from google.genai import types as genai_types
from db.connection import table
import os, time

MIN_COURSE_RELEVANCE = 0.35 # below this, document is likely off-topic for the course

try:
# Resolve BU course code from Sapling UUID
rows = table("courses").select(
"course_code", filters={"id": f"eq.{course_id}"}, limit=1
)
bu_course_id = (rows[0].get("course_code") or course_id) if rows else course_id

chunks = _chunk_text(extracted_text)
if not chunks:
return

_gclient = _genai.Client(api_key=os.getenv("GEMINI_API_KEY", ""))

def _embed_texts(texts: list[str]) -> list[list[float]]:
resp = _gclient.models.embed_content(
model="gemini-embedding-001",
contents=texts,
config=genai_types.EmbedContentConfig(output_dimensionality=768),
)
return [list(e.values) for e in resp.embeddings]

# ── Relevance gate ────────────────────────────────────────────────────
# Fetch the catalog chunk embedding for this course and compare against
# the document's first chunk. Irrelevant documents are skipped to keep
# the index clean.
catalog_rows = table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
if catalog_rows and catalog_rows[0].get("embedding"):
catalog_vec = catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))
if dot < MIN_COURSE_RELEVANCE:
Comment on lines +960 to +975

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Compare against the best catalog match, not an arbitrary catalog row.

Line 960 fetches limit=1 with no ordering, so the relevance gate can reject an on-topic upload if Supabase returns a non-representative catalog chunk. Fetch the course’s catalog embeddings and gate on the max cosine score instead.

Suggested adjustment
- catalog_rows = table("course_chunks").select(- "embedding",- filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},- limit=1,- )- if catalog_rows and catalog_rows[0].get("embedding"):- catalog_vec = catalog_rows[0]["embedding"]+ catalog_rows = table("course_chunks").select(+ "embedding",+ filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},+ )+ catalog_vecs = [r["embedding"] for r in catalog_rows if r.get("embedding")]+ if catalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
- # cosine similarity (vectors are unit-norm from the model)- dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))+ dot = max(sum(a * b for a, b in zip(doc_sample_vec, catalog_vec)) for catalog_vec in catalog_vecs)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
ifcatalog_rowsandcatalog_rows[0].get("embedding"):
catalog_vec=catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot=sum(a*bfora, binzip(doc_sample_vec, catalog_vec))
ifdot<MIN_COURSE_RELEVANCE:
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
)
catalog_vecs= [r["embedding"] forrincatalog_rowsifr.get("embedding")]
ifcatalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
dot=max(sum(a*bfora, binzip(doc_sample_vec, catalog_vec)) forcatalog_vecincatalog_vecs)
ifdot<MIN_COURSE_RELEVANCE:
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 960 - 975, The relevance check
currently uses only the first returned catalog chunk from
table("course_chunks"), which is arbitrary because the query has limit=1 without
ordering. Update the logic around the catalog_rows lookup to evaluate all
catalog embeddings for the course and compare the document sample against the
best cosine match, then apply the MIN_COURSE_RELEVANCE gate to that maximum
score. Keep the change localized near the existing sample_text, doc_sample_vec,
and dot calculation so the representative catalog selection is no longer
dependent on Supabase row order.

logger.warning(
"[RAG] doc %s skipped — relevance to %s is %.3f (< %.2f)",
doc_id, bu_course_id, dot, MIN_COURSE_RELEVANCE,
)
return

records = []
for i, chunk_text in enumerate(chunks):
raw = f"{doc_id}::{i}::{chunk_text}"
cid = hashlib.sha256(raw.encode()).hexdigest()
records.append({
"id": cid,
"course_id": bu_course_id,
"doc_id": doc_id,
"uploader_id": user_id,
"chunk_index": i,
"chunk_text": chunk_text,
"chunk_hash": cid,
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})

# Embed in batches of 50
BATCH = 50
for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota

table("course_chunks").upsert(records, on_conflict="id")
Comment on lines +994 to +1014

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Only upsert successfully embedded chunks, and keep upserts batched.

If _embed_texts fails for a batch, those records keep embedding=None but are still sent to course_chunks; additionally, all chunks are posted in one large request after embedding. That can either pollute the retrieval index or make large uploads fail at the final Supabase call. Upsert only records with embeddings per batch, matching the catalog ingest pattern.

Suggested adjustment
 for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
+ if len(vecs) != len(batch):+ raise ValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
+ table("course_chunks").upsert(batch, on_conflict="id")
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
- table("course_chunks").upsert(records, on_conflict="id")
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
table("course_chunks").upsert(records, on_conflict="id")
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
iflen(vecs) !=len(batch):
raiseValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
table("course_chunks").upsert(batch, on_conflict="id")
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 994 - 1014, The batch embedding
flow in the document ingest path still upserts every record at the end,
including chunks whose embedding stayed None after _embed_texts failed. Update
the course_chunks ingest logic to mirror the catalog ingest pattern by upserting
only the successfully embedded records inside each batch loop, and skip failed
batch records instead of carrying them forward. Keep the batching behavior in
place around the existing record processing in the documents route so large
uploads do not get sent as one final upsert.

logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
except Exception:
logger.exception("[RAG] _index_document_chunks failed for doc %s", doc_id)


def _spawn_post_roll(*tasks: tuple) -> None:
"""Fire-and-forget post-roll work for SSE / non-FastAPI-BackgroundTasks
contexts. Each tuple is (label, callable, *args). Exceptions in the
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -34,6 +34,9 @@ dist/
*.sqlite
*.sqlite3

# ---- RAG data (large files — stored in Supabase Storage) ----
backend/data/*.json

# ---- Env / secrets ----
.env
.env.*
Expand Down
210 changes: 210 additions & 0 deletions backend/data/scrape_summary.json
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,210 @@
{
"total_courses": 8720,
"total_errors": 50,
"errors": [
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-650a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-or-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-529a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-641a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-525a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-511a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-541a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-532a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-660a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-523a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-644a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-581a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pa-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-534a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-md-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-546a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-542a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-519a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-527a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-544a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-540a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-532a/",
"error": "empty response"
}
],
Comment on lines +2 to +205

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Don't commit a scrape summary that already records 50 failed course pages.

Ingestion uses this scrape output as the source corpus, so these unrecovered SDM URLs mean the initial RAG index is knowingly incomplete. Please either rerun until this is clean or make ingestion fail fast when the scrape summary reports unresolved errors.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` around lines 2 - 205, The scrape summary
currently includes unresolved failures, so do not keep this snapshot as-is.
Update the ingestion flow that reads scrape_summary.json to either rerun
scraping until the errors array is empty or make the ingest step fail fast when
total_errors is nonzero, using the summary fields total_errors and errors as the
check. Locate the validation in the scrape-summary ingestion path and ensure it
blocks indexing when any failed course URLs remain.

"elapsed_seconds": 1890,
"completed_at": "2026-06-27T01:39:55.726124+00:00",
"semester_tag": "fall_2026",
"output_file": "C:\\Users\\Jack\\Desktop\\VS Code\\sapling\\backend\\data\\bu_catalog_fall_2026.json"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

Remove the machine-local absolute path from this committed artifact.

C:\Users\Jack\... leaks workstation details and makes the summary non-portable. A repo-relative path or no path field at all would avoid noisy diffs and the local identifier exposure.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` at line 209, The committed scrape_summary
artifact currently hardcodes a machine-local absolute path in the output_file
field, which leaks workstation-specific details and breaks portability. Update
the value in scrape_summary.json to use a repo-relative path or remove the field
entirely, and ensure any code that writes this summary uses a stable path source
so future runs do not reintroduce local identifiers; check the summary
generation logic and the output_file serialization path.

}
8 changes: 8 additions & 0 deletions backend/db/connection.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -101,3 +101,11 @@ def delete(self, filters: dict) -> list:

def table(name: str) -> SupabaseTable:
return SupabaseTable(name)


def rpc(function_name: str, params: dict) -> list:
"""Call a Supabase Postgres function via /rest/v1/rpc/{function_name}."""
url = f"{REST_URL}/rpc/{function_name}"
r = _client.post(url, json=params)
r.raise_for_status()
return r.json()
119 changes: 119 additions & 0 deletions backend/routes/documents.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -785,6 +785,7 @@ async def event_stream():
("invalidate_study_guide_cache", _invalidate_study_guide_cache, user_id, course_id),
("update_course_context", update_course_context, course_id),
("check_upload_achievements", _check_upload_achievements, user_id),
("index_document_chunks", _index_document_chunks, doc_id, course_id, user_id, extracted_text, classification.category, getattr(summary, "abstract", "")),
)

yield sapling_event_to_sse(SaplingEvent(
Expand DownExpand Up@@ -898,6 +899,124 @@ def _check_upload_achievements(user_id: str) -> None:
pass


def _chunk_text(text: str, chunk_size: int = 800, overlap: int = 100) -> list[str]:
"""Split text into overlapping character-window chunks."""
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end].strip())
start += chunk_size - overlap
return [c for c in chunks if len(c) > 50] # drop near-empty tail chunks


def _index_document_chunks(
doc_id: str,
course_id: str, # Sapling UUID — resolved to BU code internally
user_id: str,
extracted_text: str,
category: str,
doc_summary: str = "",
) -> None:
"""Chunk, embed, and upsert a document into course_chunks.

Runs in a background thread via _spawn_post_roll after the document
is persisted, so it never blocks the SSE stream.
"""
import hashlib
import math
from google import genai as _genai
from google.genai import types as genai_types
from db.connection import table
import os, time

MIN_COURSE_RELEVANCE = 0.35 # below this, document is likely off-topic for the course

try:
# Resolve BU course code from Sapling UUID
rows = table("courses").select(
"course_code", filters={"id": f"eq.{course_id}"}, limit=1
)
bu_course_id = (rows[0].get("course_code") or course_id) if rows else course_id

chunks = _chunk_text(extracted_text)
if not chunks:
return

_gclient = _genai.Client(api_key=os.getenv("GEMINI_API_KEY", ""))

def _embed_texts(texts: list[str]) -> list[list[float]]:
resp = _gclient.models.embed_content(
model="gemini-embedding-001",
contents=texts,
config=genai_types.EmbedContentConfig(output_dimensionality=768),
)
return [list(e.values) for e in resp.embeddings]

# ── Relevance gate ────────────────────────────────────────────────────
# Fetch the catalog chunk embedding for this course and compare against
# the document's first chunk. Irrelevant documents are skipped to keep
# the index clean.
catalog_rows = table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
if catalog_rows and catalog_rows[0].get("embedding"):
catalog_vec = catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))
if dot < MIN_COURSE_RELEVANCE:
Comment on lines +960 to +975

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Compare against the best catalog match, not an arbitrary catalog row.

Line 960 fetches limit=1 with no ordering, so the relevance gate can reject an on-topic upload if Supabase returns a non-representative catalog chunk. Fetch the course’s catalog embeddings and gate on the max cosine score instead.

Suggested adjustment
- catalog_rows = table("course_chunks").select(- "embedding",- filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},- limit=1,- )- if catalog_rows and catalog_rows[0].get("embedding"):- catalog_vec = catalog_rows[0]["embedding"]+ catalog_rows = table("course_chunks").select(+ "embedding",+ filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},+ )+ catalog_vecs = [r["embedding"] for r in catalog_rows if r.get("embedding")]+ if catalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
- # cosine similarity (vectors are unit-norm from the model)- dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))+ dot = max(sum(a * b for a, b in zip(doc_sample_vec, catalog_vec)) for catalog_vec in catalog_vecs)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
ifcatalog_rowsandcatalog_rows[0].get("embedding"):
catalog_vec=catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot=sum(a*bfora, binzip(doc_sample_vec, catalog_vec))
ifdot<MIN_COURSE_RELEVANCE:
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
)
catalog_vecs= [r["embedding"] forrincatalog_rowsifr.get("embedding")]
ifcatalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
dot=max(sum(a*bfora, binzip(doc_sample_vec, catalog_vec)) forcatalog_vecincatalog_vecs)
ifdot<MIN_COURSE_RELEVANCE:
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 960 - 975, The relevance check
currently uses only the first returned catalog chunk from
table("course_chunks"), which is arbitrary because the query has limit=1 without
ordering. Update the logic around the catalog_rows lookup to evaluate all
catalog embeddings for the course and compare the document sample against the
best cosine match, then apply the MIN_COURSE_RELEVANCE gate to that maximum
score. Keep the change localized near the existing sample_text, doc_sample_vec,
and dot calculation so the representative catalog selection is no longer
dependent on Supabase row order.

logger.warning(
"[RAG] doc %s skipped — relevance to %s is %.3f (< %.2f)",
doc_id, bu_course_id, dot, MIN_COURSE_RELEVANCE,
)
return

records = []
for i, chunk_text in enumerate(chunks):
raw = f"{doc_id}::{i}::{chunk_text}"
cid = hashlib.sha256(raw.encode()).hexdigest()
records.append({
"id": cid,
"course_id": bu_course_id,
"doc_id": doc_id,
"uploader_id": user_id,
"chunk_index": i,
"chunk_text": chunk_text,
"chunk_hash": cid,
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})

# Embed in batches of 50
BATCH = 50
for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota

table("course_chunks").upsert(records, on_conflict="id")
Comment on lines +994 to +1014

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Only upsert successfully embedded chunks, and keep upserts batched.

If _embed_texts fails for a batch, those records keep embedding=None but are still sent to course_chunks; additionally, all chunks are posted in one large request after embedding. That can either pollute the retrieval index or make large uploads fail at the final Supabase call. Upsert only records with embeddings per batch, matching the catalog ingest pattern.

Suggested adjustment
 for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
+ if len(vecs) != len(batch):+ raise ValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
+ table("course_chunks").upsert(batch, on_conflict="id")
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
- table("course_chunks").upsert(records, on_conflict="id")
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
table("course_chunks").upsert(records, on_conflict="id")
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
iflen(vecs) !=len(batch):
raiseValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
table("course_chunks").upsert(batch, on_conflict="id")
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 994 - 1014, The batch embedding
flow in the document ingest path still upserts every record at the end,
including chunks whose embedding stayed None after _embed_texts failed. Update
the course_chunks ingest logic to mirror the catalog ingest pattern by upserting
only the successfully embedded records inside each batch loop, and skip failed
batch records instead of carrying them forward. Keep the batching behavior in
place around the existing record processing in the documents route so large
uploads do not get sent as one final upsert.

logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
except Exception:
logger.exception("[RAG] _index_document_chunks failed for doc %s", doc_id)


def _spawn_post_roll(*tasks: tuple) -> None:
"""Fire-and-forget post-roll work for SSE / non-FastAPI-BackgroundTasks
contexts. Each tuple is (label, callable, *args). Exceptions in the
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -34,6 +34,9 @@ dist/
*.sqlite
*.sqlite3

# ---- RAG data (large files — stored in Supabase Storage) ----
backend/data/*.json

# ---- Env / secrets ----
.env
.env.*
Expand Down
210 changes: 210 additions & 0 deletions backend/data/scrape_summary.json
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,210 @@
{
"total_courses": 8720,
"total_errors": 50,
"errors": [
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-650a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-or-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-529a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-641a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-525a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-511a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-541a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-532a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-660a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-523a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-644a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-581a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pa-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-534a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-md-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-546a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-542a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-519a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-527a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-544a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-540a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-532a/",
"error": "empty response"
}
],
Comment on lines +2 to +205

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Don't commit a scrape summary that already records 50 failed course pages.

Ingestion uses this scrape output as the source corpus, so these unrecovered SDM URLs mean the initial RAG index is knowingly incomplete. Please either rerun until this is clean or make ingestion fail fast when the scrape summary reports unresolved errors.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` around lines 2 - 205, The scrape summary
currently includes unresolved failures, so do not keep this snapshot as-is.
Update the ingestion flow that reads scrape_summary.json to either rerun
scraping until the errors array is empty or make the ingest step fail fast when
total_errors is nonzero, using the summary fields total_errors and errors as the
check. Locate the validation in the scrape-summary ingestion path and ensure it
blocks indexing when any failed course URLs remain.

"elapsed_seconds": 1890,
"completed_at": "2026-06-27T01:39:55.726124+00:00",
"semester_tag": "fall_2026",
"output_file": "C:\\Users\\Jack\\Desktop\\VS Code\\sapling\\backend\\data\\bu_catalog_fall_2026.json"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

Remove the machine-local absolute path from this committed artifact.

C:\Users\Jack\... leaks workstation details and makes the summary non-portable. A repo-relative path or no path field at all would avoid noisy diffs and the local identifier exposure.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` at line 209, The committed scrape_summary
artifact currently hardcodes a machine-local absolute path in the output_file
field, which leaks workstation-specific details and breaks portability. Update
the value in scrape_summary.json to use a repo-relative path or remove the field
entirely, and ensure any code that writes this summary uses a stable path source
so future runs do not reintroduce local identifiers; check the summary
generation logic and the output_file serialization path.

}
8 changes: 8 additions & 0 deletions backend/db/connection.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -101,3 +101,11 @@ def delete(self, filters: dict) -> list:

def table(name: str) -> SupabaseTable:
return SupabaseTable(name)


def rpc(function_name: str, params: dict) -> list:
"""Call a Supabase Postgres function via /rest/v1/rpc/{function_name}."""
url = f"{REST_URL}/rpc/{function_name}"
r = _client.post(url, json=params)
r.raise_for_status()
return r.json()
119 changes: 119 additions & 0 deletions backend/routes/documents.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -785,6 +785,7 @@ async def event_stream():
("invalidate_study_guide_cache", _invalidate_study_guide_cache, user_id, course_id),
("update_course_context", update_course_context, course_id),
("check_upload_achievements", _check_upload_achievements, user_id),
("index_document_chunks", _index_document_chunks, doc_id, course_id, user_id, extracted_text, classification.category, getattr(summary, "abstract", "")),
)

yield sapling_event_to_sse(SaplingEvent(
Expand DownExpand Up@@ -898,6 +899,124 @@ def _check_upload_achievements(user_id: str) -> None:
pass


def _chunk_text(text: str, chunk_size: int = 800, overlap: int = 100) -> list[str]:
"""Split text into overlapping character-window chunks."""
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end].strip())
start += chunk_size - overlap
return [c for c in chunks if len(c) > 50] # drop near-empty tail chunks


def _index_document_chunks(
doc_id: str,
course_id: str, # Sapling UUID — resolved to BU code internally
user_id: str,
extracted_text: str,
category: str,
doc_summary: str = "",
) -> None:
"""Chunk, embed, and upsert a document into course_chunks.

Runs in a background thread via _spawn_post_roll after the document
is persisted, so it never blocks the SSE stream.
"""
import hashlib
import math
from google import genai as _genai
from google.genai import types as genai_types
from db.connection import table
import os, time

MIN_COURSE_RELEVANCE = 0.35 # below this, document is likely off-topic for the course

try:
# Resolve BU course code from Sapling UUID
rows = table("courses").select(
"course_code", filters={"id": f"eq.{course_id}"}, limit=1
)
bu_course_id = (rows[0].get("course_code") or course_id) if rows else course_id

chunks = _chunk_text(extracted_text)
if not chunks:
return

_gclient = _genai.Client(api_key=os.getenv("GEMINI_API_KEY", ""))

def _embed_texts(texts: list[str]) -> list[list[float]]:
resp = _gclient.models.embed_content(
model="gemini-embedding-001",
contents=texts,
config=genai_types.EmbedContentConfig(output_dimensionality=768),
)
return [list(e.values) for e in resp.embeddings]

# ── Relevance gate ────────────────────────────────────────────────────
# Fetch the catalog chunk embedding for this course and compare against
# the document's first chunk. Irrelevant documents are skipped to keep
# the index clean.
catalog_rows = table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
if catalog_rows and catalog_rows[0].get("embedding"):
catalog_vec = catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))
if dot < MIN_COURSE_RELEVANCE:
Comment on lines +960 to +975

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Compare against the best catalog match, not an arbitrary catalog row.

Line 960 fetches limit=1 with no ordering, so the relevance gate can reject an on-topic upload if Supabase returns a non-representative catalog chunk. Fetch the course’s catalog embeddings and gate on the max cosine score instead.

Suggested adjustment
- catalog_rows = table("course_chunks").select(- "embedding",- filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},- limit=1,- )- if catalog_rows and catalog_rows[0].get("embedding"):- catalog_vec = catalog_rows[0]["embedding"]+ catalog_rows = table("course_chunks").select(+ "embedding",+ filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},+ )+ catalog_vecs = [r["embedding"] for r in catalog_rows if r.get("embedding")]+ if catalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
- # cosine similarity (vectors are unit-norm from the model)- dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))+ dot = max(sum(a * b for a, b in zip(doc_sample_vec, catalog_vec)) for catalog_vec in catalog_vecs)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
ifcatalog_rowsandcatalog_rows[0].get("embedding"):
catalog_vec=catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot=sum(a*bfora, binzip(doc_sample_vec, catalog_vec))
ifdot<MIN_COURSE_RELEVANCE:
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
)
catalog_vecs= [r["embedding"] forrincatalog_rowsifr.get("embedding")]
ifcatalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
dot=max(sum(a*bfora, binzip(doc_sample_vec, catalog_vec)) forcatalog_vecincatalog_vecs)
ifdot<MIN_COURSE_RELEVANCE:
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 960 - 975, The relevance check
currently uses only the first returned catalog chunk from
table("course_chunks"), which is arbitrary because the query has limit=1 without
ordering. Update the logic around the catalog_rows lookup to evaluate all
catalog embeddings for the course and compare the document sample against the
best cosine match, then apply the MIN_COURSE_RELEVANCE gate to that maximum
score. Keep the change localized near the existing sample_text, doc_sample_vec,
and dot calculation so the representative catalog selection is no longer
dependent on Supabase row order.

logger.warning(
"[RAG] doc %s skipped — relevance to %s is %.3f (< %.2f)",
doc_id, bu_course_id, dot, MIN_COURSE_RELEVANCE,
)
return

records = []
for i, chunk_text in enumerate(chunks):
raw = f"{doc_id}::{i}::{chunk_text}"
cid = hashlib.sha256(raw.encode()).hexdigest()
records.append({
"id": cid,
"course_id": bu_course_id,
"doc_id": doc_id,
"uploader_id": user_id,
"chunk_index": i,
"chunk_text": chunk_text,
"chunk_hash": cid,
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})

# Embed in batches of 50
BATCH = 50
for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota

table("course_chunks").upsert(records, on_conflict="id")
Comment on lines +994 to +1014

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Only upsert successfully embedded chunks, and keep upserts batched.

If _embed_texts fails for a batch, those records keep embedding=None but are still sent to course_chunks; additionally, all chunks are posted in one large request after embedding. That can either pollute the retrieval index or make large uploads fail at the final Supabase call. Upsert only records with embeddings per batch, matching the catalog ingest pattern.

Suggested adjustment
 for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
+ if len(vecs) != len(batch):+ raise ValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
+ table("course_chunks").upsert(batch, on_conflict="id")
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
- table("course_chunks").upsert(records, on_conflict="id")
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
table("course_chunks").upsert(records, on_conflict="id")
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
iflen(vecs) !=len(batch):
raiseValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
table("course_chunks").upsert(batch, on_conflict="id")
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 994 - 1014, The batch embedding
flow in the document ingest path still upserts every record at the end,
including chunks whose embedding stayed None after _embed_texts failed. Update
the course_chunks ingest logic to mirror the catalog ingest pattern by upserting
only the successfully embedded records inside each batch loop, and skip failed
batch records instead of carrying them forward. Keep the batching behavior in
place around the existing record processing in the documents route so large
uploads do not get sent as one final upsert.

logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
except Exception:
logger.exception("[RAG] _index_document_chunks failed for doc %s", doc_id)


def _spawn_post_roll(*tasks: tuple) -> None:
"""Fire-and-forget post-roll work for SSE / non-FastAPI-BackgroundTasks
contexts. Each tuple is (label, callable, *args). Exceptions in the
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -34,6 +34,9 @@ dist/
*.sqlite
*.sqlite3

# ---- RAG data (large files — stored in Supabase Storage) ----
backend/data/*.json

# ---- Env / secrets ----
.env
.env.*
Expand Down
210 changes: 210 additions & 0 deletions backend/data/scrape_summary.json
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,210 @@
{
"total_courses": 8720,
"total_errors": 50,
"errors": [
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-650a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-or-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-529a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-641a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-525a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-511a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-541a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-532a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-660a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-523a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-644a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-581a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pa-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-534a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-md-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-546a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-542a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-519a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-527a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-544a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-540a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-532a/",
"error": "empty response"
}
],
Comment on lines +2 to +205

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Don't commit a scrape summary that already records 50 failed course pages.

Ingestion uses this scrape output as the source corpus, so these unrecovered SDM URLs mean the initial RAG index is knowingly incomplete. Please either rerun until this is clean or make ingestion fail fast when the scrape summary reports unresolved errors.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` around lines 2 - 205, The scrape summary
currently includes unresolved failures, so do not keep this snapshot as-is.
Update the ingestion flow that reads scrape_summary.json to either rerun
scraping until the errors array is empty or make the ingest step fail fast when
total_errors is nonzero, using the summary fields total_errors and errors as the
check. Locate the validation in the scrape-summary ingestion path and ensure it
blocks indexing when any failed course URLs remain.

"elapsed_seconds": 1890,
"completed_at": "2026-06-27T01:39:55.726124+00:00",
"semester_tag": "fall_2026",
"output_file": "C:\\Users\\Jack\\Desktop\\VS Code\\sapling\\backend\\data\\bu_catalog_fall_2026.json"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

Remove the machine-local absolute path from this committed artifact.

C:\Users\Jack\... leaks workstation details and makes the summary non-portable. A repo-relative path or no path field at all would avoid noisy diffs and the local identifier exposure.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` at line 209, The committed scrape_summary
artifact currently hardcodes a machine-local absolute path in the output_file
field, which leaks workstation-specific details and breaks portability. Update
the value in scrape_summary.json to use a repo-relative path or remove the field
entirely, and ensure any code that writes this summary uses a stable path source
so future runs do not reintroduce local identifiers; check the summary
generation logic and the output_file serialization path.

}
8 changes: 8 additions & 0 deletions backend/db/connection.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -101,3 +101,11 @@ def delete(self, filters: dict) -> list:

def table(name: str) -> SupabaseTable:
return SupabaseTable(name)


def rpc(function_name: str, params: dict) -> list:
"""Call a Supabase Postgres function via /rest/v1/rpc/{function_name}."""
url = f"{REST_URL}/rpc/{function_name}"
r = _client.post(url, json=params)
r.raise_for_status()
return r.json()
119 changes: 119 additions & 0 deletions backend/routes/documents.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -785,6 +785,7 @@ async def event_stream():
("invalidate_study_guide_cache", _invalidate_study_guide_cache, user_id, course_id),
("update_course_context", update_course_context, course_id),
("check_upload_achievements", _check_upload_achievements, user_id),
("index_document_chunks", _index_document_chunks, doc_id, course_id, user_id, extracted_text, classification.category, getattr(summary, "abstract", "")),
)

yield sapling_event_to_sse(SaplingEvent(
Expand DownExpand Up@@ -898,6 +899,124 @@ def _check_upload_achievements(user_id: str) -> None:
pass


def _chunk_text(text: str, chunk_size: int = 800, overlap: int = 100) -> list[str]:
"""Split text into overlapping character-window chunks."""
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end].strip())
start += chunk_size - overlap
return [c for c in chunks if len(c) > 50] # drop near-empty tail chunks


def _index_document_chunks(
doc_id: str,
course_id: str, # Sapling UUID — resolved to BU code internally
user_id: str,
extracted_text: str,
category: str,
doc_summary: str = "",
) -> None:
"""Chunk, embed, and upsert a document into course_chunks.

Runs in a background thread via _spawn_post_roll after the document
is persisted, so it never blocks the SSE stream.
"""
import hashlib
import math
from google import genai as _genai
from google.genai import types as genai_types
from db.connection import table
import os, time

MIN_COURSE_RELEVANCE = 0.35 # below this, document is likely off-topic for the course

try:
# Resolve BU course code from Sapling UUID
rows = table("courses").select(
"course_code", filters={"id": f"eq.{course_id}"}, limit=1
)
bu_course_id = (rows[0].get("course_code") or course_id) if rows else course_id

chunks = _chunk_text(extracted_text)
if not chunks:
return

_gclient = _genai.Client(api_key=os.getenv("GEMINI_API_KEY", ""))

def _embed_texts(texts: list[str]) -> list[list[float]]:
resp = _gclient.models.embed_content(
model="gemini-embedding-001",
contents=texts,
config=genai_types.EmbedContentConfig(output_dimensionality=768),
)
return [list(e.values) for e in resp.embeddings]

# ── Relevance gate ────────────────────────────────────────────────────
# Fetch the catalog chunk embedding for this course and compare against
# the document's first chunk. Irrelevant documents are skipped to keep
# the index clean.
catalog_rows = table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
if catalog_rows and catalog_rows[0].get("embedding"):
catalog_vec = catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))
if dot < MIN_COURSE_RELEVANCE:
Comment on lines +960 to +975

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Compare against the best catalog match, not an arbitrary catalog row.

Line 960 fetches limit=1 with no ordering, so the relevance gate can reject an on-topic upload if Supabase returns a non-representative catalog chunk. Fetch the course’s catalog embeddings and gate on the max cosine score instead.

Suggested adjustment
- catalog_rows = table("course_chunks").select(- "embedding",- filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},- limit=1,- )- if catalog_rows and catalog_rows[0].get("embedding"):- catalog_vec = catalog_rows[0]["embedding"]+ catalog_rows = table("course_chunks").select(+ "embedding",+ filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},+ )+ catalog_vecs = [r["embedding"] for r in catalog_rows if r.get("embedding")]+ if catalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
- # cosine similarity (vectors are unit-norm from the model)- dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))+ dot = max(sum(a * b for a, b in zip(doc_sample_vec, catalog_vec)) for catalog_vec in catalog_vecs)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
ifcatalog_rowsandcatalog_rows[0].get("embedding"):
catalog_vec=catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot=sum(a*bfora, binzip(doc_sample_vec, catalog_vec))
ifdot<MIN_COURSE_RELEVANCE:
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
)
catalog_vecs= [r["embedding"] forrincatalog_rowsifr.get("embedding")]
ifcatalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
dot=max(sum(a*bfora, binzip(doc_sample_vec, catalog_vec)) forcatalog_vecincatalog_vecs)
ifdot<MIN_COURSE_RELEVANCE:
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 960 - 975, The relevance check
currently uses only the first returned catalog chunk from
table("course_chunks"), which is arbitrary because the query has limit=1 without
ordering. Update the logic around the catalog_rows lookup to evaluate all
catalog embeddings for the course and compare the document sample against the
best cosine match, then apply the MIN_COURSE_RELEVANCE gate to that maximum
score. Keep the change localized near the existing sample_text, doc_sample_vec,
and dot calculation so the representative catalog selection is no longer
dependent on Supabase row order.

logger.warning(
"[RAG] doc %s skipped — relevance to %s is %.3f (< %.2f)",
doc_id, bu_course_id, dot, MIN_COURSE_RELEVANCE,
)
return

records = []
for i, chunk_text in enumerate(chunks):
raw = f"{doc_id}::{i}::{chunk_text}"
cid = hashlib.sha256(raw.encode()).hexdigest()
records.append({
"id": cid,
"course_id": bu_course_id,
"doc_id": doc_id,
"uploader_id": user_id,
"chunk_index": i,
"chunk_text": chunk_text,
"chunk_hash": cid,
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})

# Embed in batches of 50
BATCH = 50
for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota

table("course_chunks").upsert(records, on_conflict="id")
Comment on lines +994 to +1014

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Only upsert successfully embedded chunks, and keep upserts batched.

If _embed_texts fails for a batch, those records keep embedding=None but are still sent to course_chunks; additionally, all chunks are posted in one large request after embedding. That can either pollute the retrieval index or make large uploads fail at the final Supabase call. Upsert only records with embeddings per batch, matching the catalog ingest pattern.

Suggested adjustment
 for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
+ if len(vecs) != len(batch):+ raise ValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
+ table("course_chunks").upsert(batch, on_conflict="id")
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
- table("course_chunks").upsert(records, on_conflict="id")
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
table("course_chunks").upsert(records, on_conflict="id")
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
iflen(vecs) !=len(batch):
raiseValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
table("course_chunks").upsert(batch, on_conflict="id")
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 994 - 1014, The batch embedding
flow in the document ingest path still upserts every record at the end,
including chunks whose embedding stayed None after _embed_texts failed. Update
the course_chunks ingest logic to mirror the catalog ingest pattern by upserting
only the successfully embedded records inside each batch loop, and skip failed
batch records instead of carrying them forward. Keep the batching behavior in
place around the existing record processing in the documents route so large
uploads do not get sent as one final upsert.

logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
except Exception:
logger.exception("[RAG] _index_document_chunks failed for doc %s", doc_id)


def _spawn_post_roll(*tasks: tuple) -> None:
"""Fire-and-forget post-roll work for SSE / non-FastAPI-BackgroundTasks
contexts. Each tuple is (label, callable, *args). Exceptions in the
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -34,6 +34,9 @@ dist/
*.sqlite
*.sqlite3

# ---- RAG data (large files — stored in Supabase Storage) ----
backend/data/*.json

# ---- Env / secrets ----
.env
.env.*
Expand Down
210 changes: 210 additions & 0 deletions backend/data/scrape_summary.json
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,210 @@
{
"total_courses": 8720,
"total_errors": 50,
"errors": [
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-650a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-or-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-529a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-641a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-525a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-511a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-541a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-532a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-660a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-523a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-644a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-581a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pa-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-534a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-md-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-546a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-542a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-519a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-527a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-544a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-540a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-532a/",
"error": "empty response"
}
],
Comment on lines +2 to +205

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Don't commit a scrape summary that already records 50 failed course pages.

Ingestion uses this scrape output as the source corpus, so these unrecovered SDM URLs mean the initial RAG index is knowingly incomplete. Please either rerun until this is clean or make ingestion fail fast when the scrape summary reports unresolved errors.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` around lines 2 - 205, The scrape summary
currently includes unresolved failures, so do not keep this snapshot as-is.
Update the ingestion flow that reads scrape_summary.json to either rerun
scraping until the errors array is empty or make the ingest step fail fast when
total_errors is nonzero, using the summary fields total_errors and errors as the
check. Locate the validation in the scrape-summary ingestion path and ensure it
blocks indexing when any failed course URLs remain.

"elapsed_seconds": 1890,
"completed_at": "2026-06-27T01:39:55.726124+00:00",
"semester_tag": "fall_2026",
"output_file": "C:\\Users\\Jack\\Desktop\\VS Code\\sapling\\backend\\data\\bu_catalog_fall_2026.json"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

Remove the machine-local absolute path from this committed artifact.

C:\Users\Jack\... leaks workstation details and makes the summary non-portable. A repo-relative path or no path field at all would avoid noisy diffs and the local identifier exposure.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` at line 209, The committed scrape_summary
artifact currently hardcodes a machine-local absolute path in the output_file
field, which leaks workstation-specific details and breaks portability. Update
the value in scrape_summary.json to use a repo-relative path or remove the field
entirely, and ensure any code that writes this summary uses a stable path source
so future runs do not reintroduce local identifiers; check the summary
generation logic and the output_file serialization path.

}
8 changes: 8 additions & 0 deletions backend/db/connection.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -101,3 +101,11 @@ def delete(self, filters: dict) -> list:

def table(name: str) -> SupabaseTable:
return SupabaseTable(name)


def rpc(function_name: str, params: dict) -> list:
"""Call a Supabase Postgres function via /rest/v1/rpc/{function_name}."""
url = f"{REST_URL}/rpc/{function_name}"
r = _client.post(url, json=params)
r.raise_for_status()
return r.json()
119 changes: 119 additions & 0 deletions backend/routes/documents.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -785,6 +785,7 @@ async def event_stream():
("invalidate_study_guide_cache", _invalidate_study_guide_cache, user_id, course_id),
("update_course_context", update_course_context, course_id),
("check_upload_achievements", _check_upload_achievements, user_id),
("index_document_chunks", _index_document_chunks, doc_id, course_id, user_id, extracted_text, classification.category, getattr(summary, "abstract", "")),
)

yield sapling_event_to_sse(SaplingEvent(
Expand DownExpand Up@@ -898,6 +899,124 @@ def _check_upload_achievements(user_id: str) -> None:
pass


def _chunk_text(text: str, chunk_size: int = 800, overlap: int = 100) -> list[str]:
"""Split text into overlapping character-window chunks."""
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end].strip())
start += chunk_size - overlap
return [c for c in chunks if len(c) > 50] # drop near-empty tail chunks


def _index_document_chunks(
doc_id: str,
course_id: str, # Sapling UUID — resolved to BU code internally
user_id: str,
extracted_text: str,
category: str,
doc_summary: str = "",
) -> None:
"""Chunk, embed, and upsert a document into course_chunks.

Runs in a background thread via _spawn_post_roll after the document
is persisted, so it never blocks the SSE stream.
"""
import hashlib
import math
from google import genai as _genai
from google.genai import types as genai_types
from db.connection import table
import os, time

MIN_COURSE_RELEVANCE = 0.35 # below this, document is likely off-topic for the course

try:
# Resolve BU course code from Sapling UUID
rows = table("courses").select(
"course_code", filters={"id": f"eq.{course_id}"}, limit=1
)
bu_course_id = (rows[0].get("course_code") or course_id) if rows else course_id

chunks = _chunk_text(extracted_text)
if not chunks:
return

_gclient = _genai.Client(api_key=os.getenv("GEMINI_API_KEY", ""))

def _embed_texts(texts: list[str]) -> list[list[float]]:
resp = _gclient.models.embed_content(
model="gemini-embedding-001",
contents=texts,
config=genai_types.EmbedContentConfig(output_dimensionality=768),
)
return [list(e.values) for e in resp.embeddings]

# ── Relevance gate ────────────────────────────────────────────────────
# Fetch the catalog chunk embedding for this course and compare against
# the document's first chunk. Irrelevant documents are skipped to keep
# the index clean.
catalog_rows = table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
if catalog_rows and catalog_rows[0].get("embedding"):
catalog_vec = catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))
if dot < MIN_COURSE_RELEVANCE:
Comment on lines +960 to +975

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Compare against the best catalog match, not an arbitrary catalog row.

Line 960 fetches limit=1 with no ordering, so the relevance gate can reject an on-topic upload if Supabase returns a non-representative catalog chunk. Fetch the course’s catalog embeddings and gate on the max cosine score instead.

Suggested adjustment
- catalog_rows = table("course_chunks").select(- "embedding",- filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},- limit=1,- )- if catalog_rows and catalog_rows[0].get("embedding"):- catalog_vec = catalog_rows[0]["embedding"]+ catalog_rows = table("course_chunks").select(+ "embedding",+ filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},+ )+ catalog_vecs = [r["embedding"] for r in catalog_rows if r.get("embedding")]+ if catalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
- # cosine similarity (vectors are unit-norm from the model)- dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))+ dot = max(sum(a * b for a, b in zip(doc_sample_vec, catalog_vec)) for catalog_vec in catalog_vecs)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
ifcatalog_rowsandcatalog_rows[0].get("embedding"):
catalog_vec=catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot=sum(a*bfora, binzip(doc_sample_vec, catalog_vec))
ifdot<MIN_COURSE_RELEVANCE:
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
)
catalog_vecs= [r["embedding"] forrincatalog_rowsifr.get("embedding")]
ifcatalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
dot=max(sum(a*bfora, binzip(doc_sample_vec, catalog_vec)) forcatalog_vecincatalog_vecs)
ifdot<MIN_COURSE_RELEVANCE:
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 960 - 975, The relevance check
currently uses only the first returned catalog chunk from
table("course_chunks"), which is arbitrary because the query has limit=1 without
ordering. Update the logic around the catalog_rows lookup to evaluate all
catalog embeddings for the course and compare the document sample against the
best cosine match, then apply the MIN_COURSE_RELEVANCE gate to that maximum
score. Keep the change localized near the existing sample_text, doc_sample_vec,
and dot calculation so the representative catalog selection is no longer
dependent on Supabase row order.

logger.warning(
"[RAG] doc %s skipped — relevance to %s is %.3f (< %.2f)",
doc_id, bu_course_id, dot, MIN_COURSE_RELEVANCE,
)
return

records = []
for i, chunk_text in enumerate(chunks):
raw = f"{doc_id}::{i}::{chunk_text}"
cid = hashlib.sha256(raw.encode()).hexdigest()
records.append({
"id": cid,
"course_id": bu_course_id,
"doc_id": doc_id,
"uploader_id": user_id,
"chunk_index": i,
"chunk_text": chunk_text,
"chunk_hash": cid,
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})

# Embed in batches of 50
BATCH = 50
for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota

table("course_chunks").upsert(records, on_conflict="id")
Comment on lines +994 to +1014

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Only upsert successfully embedded chunks, and keep upserts batched.

If _embed_texts fails for a batch, those records keep embedding=None but are still sent to course_chunks; additionally, all chunks are posted in one large request after embedding. That can either pollute the retrieval index or make large uploads fail at the final Supabase call. Upsert only records with embeddings per batch, matching the catalog ingest pattern.

Suggested adjustment
 for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
+ if len(vecs) != len(batch):+ raise ValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
+ table("course_chunks").upsert(batch, on_conflict="id")
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
- table("course_chunks").upsert(records, on_conflict="id")
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
table("course_chunks").upsert(records, on_conflict="id")
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
iflen(vecs) !=len(batch):
raiseValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
table("course_chunks").upsert(batch, on_conflict="id")
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 994 - 1014, The batch embedding
flow in the document ingest path still upserts every record at the end,
including chunks whose embedding stayed None after _embed_texts failed. Update
the course_chunks ingest logic to mirror the catalog ingest pattern by upserting
only the successfully embedded records inside each batch loop, and skip failed
batch records instead of carrying them forward. Keep the batching behavior in
place around the existing record processing in the documents route so large
uploads do not get sent as one final upsert.

logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
except Exception:
logger.exception("[RAG] _index_document_chunks failed for doc %s", doc_id)


def _spawn_post_roll(*tasks: tuple) -> None:
"""Fire-and-forget post-roll work for SSE / non-FastAPI-BackgroundTasks
contexts. Each tuple is (label, callable, *args). Exceptions in the
Expand Down
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line numberDiff line numberDiff line change
Expand Up@@ -34,6 +34,9 @@ dist/
*.sqlite
*.sqlite3

# ---- RAG data (large files — stored in Supabase Storage) ----
backend/data/*.json

# ---- Env / secrets ----
.env
.env.*
Expand Down
210 changes: 210 additions & 0 deletions backend/data/scrape_summary.json
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,210 @@
{
"total_courses": 8720,
"total_errors": 50,
"errors": [
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-650a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-or-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-529a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-512a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-641a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-525a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pd-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-511a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-522a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-541a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-532a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-660a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-523a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-od-644a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-581a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pa-530a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-534a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-md-531a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-546a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-542a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-en-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-519a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-pe-520a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-527a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-521a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-642a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-os-640a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-544a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-gd-540a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-ph-524a/",
"error": "empty response"
},
{
"url": "https://www.bu.edu/academics/sdm/courses/sdm-rs-532a/",
"error": "empty response"
}
],
Comment on lines +2 to +205

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Don't commit a scrape summary that already records 50 failed course pages.

Ingestion uses this scrape output as the source corpus, so these unrecovered SDM URLs mean the initial RAG index is knowingly incomplete. Please either rerun until this is clean or make ingestion fail fast when the scrape summary reports unresolved errors.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` around lines 2 - 205, The scrape summary
currently includes unresolved failures, so do not keep this snapshot as-is.
Update the ingestion flow that reads scrape_summary.json to either rerun
scraping until the errors array is empty or make the ingest step fail fast when
total_errors is nonzero, using the summary fields total_errors and errors as the
check. Locate the validation in the scrape-summary ingestion path and ensure it
blocks indexing when any failed course URLs remain.

"elapsed_seconds": 1890,
"completed_at": "2026-06-27T01:39:55.726124+00:00",
"semester_tag": "fall_2026",
"output_file": "C:\\Users\\Jack\\Desktop\\VS Code\\sapling\\backend\\data\\bu_catalog_fall_2026.json"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

Remove the machine-local absolute path from this committed artifact.

C:\Users\Jack\... leaks workstation details and makes the summary non-portable. A repo-relative path or no path field at all would avoid noisy diffs and the local identifier exposure.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/data/scrape_summary.json` at line 209, The committed scrape_summary
artifact currently hardcodes a machine-local absolute path in the output_file
field, which leaks workstation-specific details and breaks portability. Update
the value in scrape_summary.json to use a repo-relative path or remove the field
entirely, and ensure any code that writes this summary uses a stable path source
so future runs do not reintroduce local identifiers; check the summary
generation logic and the output_file serialization path.

}
8 changes: 8 additions & 0 deletions backend/db/connection.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -101,3 +101,11 @@ def delete(self, filters: dict) -> list:

def table(name: str) -> SupabaseTable:
return SupabaseTable(name)


def rpc(function_name: str, params: dict) -> list:
"""Call a Supabase Postgres function via /rest/v1/rpc/{function_name}."""
url = f"{REST_URL}/rpc/{function_name}"
r = _client.post(url, json=params)
r.raise_for_status()
return r.json()
119 changes: 119 additions & 0 deletions backend/routes/documents.py
Original file line numberDiff line numberDiff line change
Expand Up@@ -785,6 +785,7 @@ async def event_stream():
("invalidate_study_guide_cache", _invalidate_study_guide_cache, user_id, course_id),
("update_course_context", update_course_context, course_id),
("check_upload_achievements", _check_upload_achievements, user_id),
("index_document_chunks", _index_document_chunks, doc_id, course_id, user_id, extracted_text, classification.category, getattr(summary, "abstract", "")),
)

yield sapling_event_to_sse(SaplingEvent(
Expand DownExpand Up@@ -898,6 +899,124 @@ def _check_upload_achievements(user_id: str) -> None:
pass


def _chunk_text(text: str, chunk_size: int = 800, overlap: int = 100) -> list[str]:
"""Split text into overlapping character-window chunks."""
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end].strip())
start += chunk_size - overlap
return [c for c in chunks if len(c) > 50] # drop near-empty tail chunks


def _index_document_chunks(
doc_id: str,
course_id: str, # Sapling UUID — resolved to BU code internally
user_id: str,
extracted_text: str,
category: str,
doc_summary: str = "",
) -> None:
"""Chunk, embed, and upsert a document into course_chunks.

Runs in a background thread via _spawn_post_roll after the document
is persisted, so it never blocks the SSE stream.
"""
import hashlib
import math
from google import genai as _genai
from google.genai import types as genai_types
from db.connection import table
import os, time

MIN_COURSE_RELEVANCE = 0.35 # below this, document is likely off-topic for the course

try:
# Resolve BU course code from Sapling UUID
rows = table("courses").select(
"course_code", filters={"id": f"eq.{course_id}"}, limit=1
)
bu_course_id = (rows[0].get("course_code") or course_id) if rows else course_id

chunks = _chunk_text(extracted_text)
if not chunks:
return

_gclient = _genai.Client(api_key=os.getenv("GEMINI_API_KEY", ""))

def _embed_texts(texts: list[str]) -> list[list[float]]:
resp = _gclient.models.embed_content(
model="gemini-embedding-001",
contents=texts,
config=genai_types.EmbedContentConfig(output_dimensionality=768),
)
return [list(e.values) for e in resp.embeddings]

# ── Relevance gate ────────────────────────────────────────────────────
# Fetch the catalog chunk embedding for this course and compare against
# the document's first chunk. Irrelevant documents are skipped to keep
# the index clean.
catalog_rows = table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
if catalog_rows and catalog_rows[0].get("embedding"):
catalog_vec = catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))
if dot < MIN_COURSE_RELEVANCE:
Comment on lines +960 to +975

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Compare against the best catalog match, not an arbitrary catalog row.

Line 960 fetches limit=1 with no ordering, so the relevance gate can reject an on-topic upload if Supabase returns a non-representative catalog chunk. Fetch the course’s catalog embeddings and gate on the max cosine score instead.

Suggested adjustment
- catalog_rows = table("course_chunks").select(- "embedding",- filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},- limit=1,- )- if catalog_rows and catalog_rows[0].get("embedding"):- catalog_vec = catalog_rows[0]["embedding"]+ catalog_rows = table("course_chunks").select(+ "embedding",+ filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},+ )+ catalog_vecs = [r["embedding"] for r in catalog_rows if r.get("embedding")]+ if catalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text = doc_summary or chunks[0]
doc_sample_vec = _embed_texts([sample_text])[0]
time.sleep(1.5)
- # cosine similarity (vectors are unit-norm from the model)- dot = sum(a * b for a, b in zip(doc_sample_vec, catalog_vec))+ dot = max(sum(a * b for a, b in zip(doc_sample_vec, catalog_vec)) for catalog_vec in catalog_vecs)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
limit=1,
)
ifcatalog_rowsandcatalog_rows[0].get("embedding"):
catalog_vec=catalog_rows[0]["embedding"]
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
# cosine similarity (vectors are unit-norm from the model)
dot=sum(a*bfora, binzip(doc_sample_vec, catalog_vec))
ifdot<MIN_COURSE_RELEVANCE:
catalog_rows=table("course_chunks").select(
"embedding",
filters={"course_id": f"eq.{bu_course_id}", "category": "eq.catalog"},
)
catalog_vecs= [r["embedding"] forrincatalog_rowsifr.get("embedding")]
ifcatalog_vecs:
# Use the AI-generated summary as the document representative —
# it's more reliable than raw first-chunk text (avoids cover pages,
# tables of contents, and boilerplate skewing the score).
sample_text=doc_summaryorchunks[0]
doc_sample_vec=_embed_texts([sample_text])[0]
time.sleep(1.5)
dot=max(sum(a*bfora, binzip(doc_sample_vec, catalog_vec)) forcatalog_vecincatalog_vecs)
ifdot<MIN_COURSE_RELEVANCE:
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 960 - 975, The relevance check
currently uses only the first returned catalog chunk from
table("course_chunks"), which is arbitrary because the query has limit=1 without
ordering. Update the logic around the catalog_rows lookup to evaluate all
catalog embeddings for the course and compare the document sample against the
best cosine match, then apply the MIN_COURSE_RELEVANCE gate to that maximum
score. Keep the change localized near the existing sample_text, doc_sample_vec,
and dot calculation so the representative catalog selection is no longer
dependent on Supabase row order.

logger.warning(
"[RAG] doc %s skipped — relevance to %s is %.3f (< %.2f)",
doc_id, bu_course_id, dot, MIN_COURSE_RELEVANCE,
)
return

records = []
for i, chunk_text in enumerate(chunks):
raw = f"{doc_id}::{i}::{chunk_text}"
cid = hashlib.sha256(raw.encode()).hexdigest()
records.append({
"id": cid,
"course_id": bu_course_id,
"doc_id": doc_id,
"uploader_id": user_id,
"chunk_index": i,
"chunk_text": chunk_text,
"chunk_hash": cid,
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})

# Embed in batches of 50
BATCH = 50
for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota

table("course_chunks").upsert(records, on_conflict="id")
Comment on lines +994 to +1014

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Only upsert successfully embedded chunks, and keep upserts batched.

If _embed_texts fails for a batch, those records keep embedding=None but are still sent to course_chunks; additionally, all chunks are posted in one large request after embedding. That can either pollute the retrieval index or make large uploads fail at the final Supabase call. Upsert only records with embeddings per batch, matching the catalog ingest pattern.

Suggested adjustment
 for i in range(0, len(records), BATCH):
batch = records[i : i + BATCH]
texts = [r["chunk_text"] for r in batch]
try:
vecs = _embed_texts(texts)
+ if len(vecs) != len(batch):+ raise ValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
for rec, vec in zip(batch, vecs):
rec["embedding"] = vec
+ table("course_chunks").upsert(batch, on_conflict="id")
except Exception as e:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
- table("course_chunks").upsert(records, on_conflict="id")
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
table("course_chunks").upsert(records, on_conflict="id")
"embedding": None,
"category": category,
"semester": "current",
"section_id": None,
"school": "",
})
# Embed in batches of 50
BATCH=50
foriinrange(0, len(records), BATCH):
batch=records[i : i+BATCH]
texts= [r["chunk_text"] forrinbatch]
try:
vecs=_embed_texts(texts)
iflen(vecs) !=len(batch):
raiseValueError(f"embedding count mismatch: {len(vecs)} for {len(batch)} chunks")
forrec, vecinzip(batch, vecs):
rec["embedding"] =vec
table("course_chunks").upsert(batch, on_conflict="id")
exceptExceptionase:
logger.warning("[RAG] embed failed for doc %s batch %d: %s", doc_id, i, e)
time.sleep(1.5) # stay under 3000 req/min quota
logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@backend/routes/documents.py` around lines 994 - 1014, The batch embedding
flow in the document ingest path still upserts every record at the end,
including chunks whose embedding stayed None after _embed_texts failed. Update
the course_chunks ingest logic to mirror the catalog ingest pattern by upserting
only the successfully embedded records inside each batch loop, and skip failed
batch records instead of carrying them forward. Keep the batching behavior in
place around the existing record processing in the documents route so large
uploads do not get sent as one final upsert.

logger.info("[RAG] indexed %d chunks for doc %s", len(records), doc_id)
except Exception:
logger.exception("[RAG] _index_document_chunks failed for doc %s", doc_id)


def _spawn_post_roll(*tasks: tuple) -> None:
"""Fire-and-forget post-roll work for SSE / non-FastAPI-BackgroundTasks
contexts. Each tuple is (label, callable, *args). Exceptions in the
Expand Down
Loading